发表机构
Tsinghua University; Fudan University; University of Edinburgh(清华大学; 复旦大学; 爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DAFI智能体,通过主动收集证据和组件反馈,实现SAE特征的输入侧、输出侧及功能解释,显著提升解释分数和效率,并提炼技能提高泛化能力。
AI 中文摘要
稀疏自编码器(SAEs)是机制可解释性的重要工具,但解释其众多特征仍然具有挑战性。现有方法分别刻画输入侧激活模式和输出侧干预效应,却往往使它们之间的功能联系隐而不显,而输入侧证据收集通常依赖于代价高昂的大规模语料库扫描。我们引入功能解释,将SAE特征刻画为从激活输入语义到干预下输出效应的映射,并提出双端智能体特征解释(DAFI),一个通过组件特定反馈主动收集证据并细化输入侧、输出侧和功能解释的智能体。其短上下文令牌探测使得无需全语料库扫描即可按需收集激活证据。在GemmaScope上,DAFI在输入分数上比SAGE提高13.1个百分点,在输出分数上比Token Change提高38.9个百分点,同时比通用编码智能体显著更节省令牌。从成功细化中提炼的技能将保留集的联合通过率从58.0%提高到92.0%,并在迁移到新的模型-SAE设置时提高了解释质量和效率。在具有可靠端点解释的特征中,70.7%表现出非等价的输入和输出语义。在AxBench上,DAFI还改进了相对于输出分数过滤的转向特征选择。代码可在该https URL获取。
英文摘要
Sparse autoencoders (SAEs) are an important tool for mechanistic interpretability, but interpreting their many features remains challenging. Existing methods characterize input-side activation patterns and output-side intervention effects, yet often leave their functional connection implicit, while input-side evidence collection typically relies on costly large-corpus scans. We introduce functional interpretation, which characterizes an SAE feature as a mapping from its activating input semantics to its output effects under intervention, and present Dual-End Agentic Feature Interpretation (DAFI), an agent that actively gathers evidence and refines input-side, output-side, and functional interpretations through component-specific feedback. Its short-context token probing enables on-demand activation evidence collection without a full corpus scan. On GemmaScope, DAFI improves Input score by 13.1 percentage points over SAGE and Output score by 38.9 points over Token Change, while being substantially more token-efficient than a general-purpose coding agent. Skills distilled from successful refinements raise the held-out joint pass rate from 58.0% to 92.0% and improve both interpretation quality and efficiency when transferred to a new model-SAE setting. Across features with reliable endpoint interpretations, 70.7% exhibit non-equivalent input and output semantics. On AxBench, DAFI also improves steering-feature selection over output-score filtering. Code is available at https://github.com/THUAIS-Lab/DAFI.
Comments25 pages