arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09113cs.AIcs.CLcs.LG

SAEScientist-Bench:AI智能体能否进行自主SAE可解释性研究?

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

Yuqiao Tan, Shizhu He, Jun Zhao, Kang Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出SAEScientist-Bench基准,评估AI智能体利用稀疏自编码器进行自主机制可解释性研究的能力,实验显示前沿智能体具备真实发现能力但落后于专家,尤其在因果引导方面。研究确立了实验性模型理解作为自主AI研发的可衡量能力。

中文摘要 AI 辅助

虽然关于递归自我改进(RSI)的研究主要集中于自动化模型训练流程,但可靠的自主开发需要一个缺失的支柱:事后监控与审计,以理解模型所学内容并确保安全对齐。机制可解释性工具对于弥合这一差距至关重要,其中稀疏自编码器(SAEs)通过分离可解释特征以进行模型检查和引导,成为基石。在本文中,我们引入了SAEScientist-Bench来评估AI智能体是否能够作为科学家,利用SAE工具进行自主机制发现。给定一个目标概念,智能体设计对比探针,并在Gemma-2-9B-IT的Gemma Scope字典中导航131K+特征,以发现最优特征,并与锚定在Neuronpedia上的精选专家参考特征在激活排名、对比文本上的概念选择性以及因果引导方面进行评估。在10种智能体配置和20项任务中,前沿智能体展现出真正的发现能力并在不同评估维度上领先,但仍远落后于专家基线,在将目标概念与对比控制分离方面接近专家水平,而在因果生成引导方面则大幅滞后。进一步分析表明,尽管智能体能够设计对比以排除虚假候选,但它们经常误解实验测量结果。这些结果将实验性模型理解确立为闭环自主AI研发的一项可衡量能力。我们的代码可在以下网址获取:此https URL。

英文摘要

While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this gap, among which Sparse Autoencoders (SAEs) serve as a cornerstone by isolating interpretable features for model inspection and steering. In this paper, we introduce SAEScientist-Bench to evaluate whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery. Given a target concept, an agent designs contrastive probes and navigates a Gemma Scope dictionary of 131K+ features in Gemma-2-9B-IT to discover the optimal feature, evaluated against curated expert reference features anchored on Neuronpedia across activation rank, concept selectivity on contrastive texts, and causal steering. Across 10 agent configurations and 20 tasks, frontier agents demonstrate genuine discovery capabilities and lead different evaluation dimensions, but remain well behind the expert baseline, approaching expert levels on separating target concepts from contrastive controls while lagging substantially in causal generation steering. Further analysis reveals that although agents can design contrasts to rule out spurious candidates, they frequently misinterpret experimental measurements. These results establish experimental model understanding as a measurable capability for closed-loop autonomous AI R&D. Our code is available at https://github.com/Trae1ounG/SAEScientist.

发表机构

  • Institute of Automation, CAS(中国科学院自动化研究所)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑