SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? Paper • 2609.09113 • Published 6 days ago • 20
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Paper • 2608.02499 • Published Aug 3 • 24