arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

主动预算会损害灵敏度:诊断与修复TopK稀疏自编码器的可靠性

Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability

Zhenting Huang, Bo Jiang, Junnan Liu, Zhixing Tan, Qianren Mao

arXiv 2609.37857首次发表:更新:

发表机构

Beihang University; Monash University; Zhongguancun Laboratory(北京航空航天大学; 莫纳什大学; 中关村实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对宽TopK稀疏自编码器在语义变化下稀有特征灵敏度下降的问题,本文通过因子实验定位主动预算k为根因,提出基于主动边际的成对排名稳定化方法,提升稀有特征灵敏度8.83个百分点,同时保持重建与覆盖率。

AI 中文摘要

稀疏自编码器(SAEs)正日益扩展到更宽的字典,以从大型语言模型的激活中恢复细粒度结构。然而,只有当特征在相同含义以不同表面形式表达时仍保持稳定的分析单元,该特征才对解释有用。我们通过特征灵敏度研究TopK SAEs的这一可靠性问题。实验表明,扩展规模选择性地降低了稀有特征的灵敏度,而常见特征保持相对稳定。一个受控的宽度×k因子实验确定了主动预算k是根本原因:退化源于选择边界,而非仅字典宽度。我们将此失败归因于TopK选择的几何特性。主动边际,即到截止点的距离,无需阈值即可预测特征丢失。在此边际诊断的指导下,我们引入了成对排名稳定化方法。我们的方法针对截止点的排序失败,将稀有特征灵敏度提高了8.83个百分点,同时保持重建和活跃特征覆盖率接近基线。总体而言,我们的结果表明,宽TopK SAEs不仅应通过重建、稀疏性和特征数量来评估,还应通过语义变化下的特征可靠性和边界几何来评估,以实现稳定的可解释性。

英文摘要

Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. However, a feature is useful for interpretation only if it remains a stable unit of analysis when the same meaning is expressed in different surface forms. We study this reliability question for TopK SAEs via feature sensitivity. Experiments demonstrate that scaling selectively reduces the sensitivity of rare features, while common features remain comparatively stable. A controlled width\(\times k\) factorial experiment identifies the active budget k as the root cause: the degradation arises from the selection boundary rather than dictionary width alone. We attribute this failure to the geometry of TopK selection. The active margin, the distance to the cutoff, predicts feature loss without thresholds. Guided by this margin diagnosis, we introduce pairwise rank stabilization. Our method targets ordering failures at the cutoff and improves rare-feature sensitivity by \(8.83\) percentage points, while keeping reconstruction and alive-feature coverage near the baseline. Overall, our results suggest that wide TopK SAEs should be evaluated not only by reconstruction, sparsity, and feature count, but also by feature reliability under semantic variation and boundary geometry for stable interpretability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑