ZeroHour
Story · 1 source · 1 articlefirst updated ()1

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

infoAI researchimportance 55
What's new: Initial merged summary (no prior story). Combines the Hugging Face daily papers entry (2026-09-07) with the arXiv listing (2026-09-08); the two sources agree on all core claims (benchmark scope, 131K+ feature search space, 10 configurations, 20 tasks, contrastive-separation strength vs. causal-steering weakness, measurement misinterpretation). The arXiv version contributes the scoring methodology…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

SAEScientist-Bench tests whether AI agents can autonomously conduct SAE interpretability research on Gemma-2-9B-IT; across 10 agent configurations and 20 tasks, frontier agents show genuine discovery capability and approach expert levels on contrastive…

SAEScientist-Bench evaluates whether AI agents can act as autonomous scientists conducting SAE interpretability research. Agents must design contrastive probes and search the Gemma Scope dictionary of over 131,000 features in Gemma-2-9B-IT, scored against expert-curated reference features on Neuronpedia via activation rank, concept selectivity, and causal steering. Across 10 agent configurations and 20 tasks, frontier agents demonstrated genuine discovery capability and approached expert levels at separating target concepts from controls, but remained well behind expert reference features overall, lagging most on causal steering. Agents frequently misinterpreted experimental measurements (such as steering results) even when their contrast designs were effective. The authors frame the benchmark as establishing experimental model understanding as a measurable capability for closed-loop autonomous AI R&D and post-hoc monitoring and auditing for safe recursive self-improvement. Code has been publicly released on GitHub (per the Hugging Face listing).

  • Benchmark: SAEScientist-Bench, targeting autonomous SAE interpretability research by AI agents.
  • Search space: the Gemma Scope dictionary of 131K+ features in Gemma-2-9B-IT.
  • Scale: 10 agent configurations evaluated across 20 tasks.
  • Scoring (per the arXiv report): against expert-curated reference features on Neuronpedia via activation rank, concept selectivity, and causal steering.
  • Headline result: frontier agents show genuine discovery capability and approach expert level on contrastive separation, but lag substantially on causal steering; agents frequently misinterpret experimental measurements even when designing…
  • Code publicly released on GitHub (per the Hugging Face listing).
  • Stated motivation (per the arXiv report): make experimental model understanding a measurable capability for closed-loop autonomous AI R&D and post-hoc monitoring/auditing for safe recursive self-improvement.
  • Timeline: appeared on Hugging Face daily papers 2026-09-07 and in the arXiv cs.AI/cs.LG/cs.CL listing 2026-09-08.

Coverage timeline

  1. · 8d ago
    Hugging Face daily papers· 55
    SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

    SAEScientist-Bench evaluates whether AI agents can autonomously conduct SAE interpretability research in Gemma-2-9B-IT, finding frontier agents trail expert baselines.