发表机构
ETH Zürich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究稀疏自动编码器可解释性得分的跨论文比较,通过多指标、多模型及方法变化轴的实验发现方法方差主导得分方差,排名不稳定,基于得分的比较可能反映管道差异,贡献评估方法以推动可解释性研究。
AI 中文摘要
稀疏自动编码器(SAE)可解释性的跨论文比较通常依赖于自动可解释性得分。在这个评估管道中,一个语言模型(LM)解释每个特征,另一个LM对解释进行评分。为使这些比较有意义,得分必须反映特征的稳定属性而非评估管道的混杂因素。通过对四个指标(模拟、检测、模糊测试、纯度)、两个模型(Pythia - 160M、Apertus - 8B)以及方法变化的四个轴进行系统实验,发现该假设不成立。具体而言,方法方差总体超过架构方差;各指标有不同不稳定特征;top - k特征排名在语料库和绘制条件下不一致。这些发现表明基于自动可解释性得分的跨论文比较可能反映管道差异而非架构差异,对SAE效用的争论有影响。更广泛地说,不可靠评估阻碍可解释性研究进展。为支持评估,贡献了方差分解方法、稳定性检查和最低报告清单。
英文摘要
Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a language model (LM) explains each feature, and another LM scores the explanation. For these comparisons to be meaningful, scores must reflect stable properties of the features rather than confounding aspects of the evaluation pipeline. Through systematic experiments across four metrics (simulation, detection, fuzzing, purity), two models (Pythia-160M, Apertus-8B), and four axes of methodological variation, we show that this assumption does not hold. Specifically, we find that R1) methodological variance collectively exceeds architectural variance across all metrics and tested models; R2) each metric exhibits a distinct instability profile, with detection being the most stable and fuzzing unreliable across all conditions; R3) top-k feature rankings do not stay consistent across corpus and draw conditions, masking per-feature instability behind stable mean scores; a failure that cannot be detected by monitoring explanation similarity alone. These findings suggest that cross-paper comparisons based on autointerpretability scores may reflect pipeline differences rather than architectural differences, with implications for the ongoing debate on SAE utility. More broadly, unreliable evaluation slows progress in interpretability research at a time when reliable tools for understanding AI systems are needed. To support evaluation, we contribute a variance decomposition approach, a Stability Check, and a Minimum Reporting Checklist.