arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单令牌稀疏自动编码器特征在因果关系上是必要的吗?层深度和SAE家族效应

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

Seonglae Cho, Zekun Wu, Kleyton Da Costa, Rishi Kalra, Ilham Wicaksono, Adriano Koshiyama

arXiv 2607.20596首次发表:更新:

发表机构

Holistic AI; University College London(整体人工智能公司; 伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究SAE特征因果作用在不同家族是否稳定,通过分析六个模型和三个SAE家族的390万个特征,发现单令牌特征聚类紧密且集中在早期层,跨家族因果有差异,其可解释性受训练方法影响,非仅激活函数或规模。

AI 中文摘要

稀疏自动编码器(SAE)特征用于解释和引导大语言模型,但一个特征的因果作用在SAE家族中是否稳定仍未得到检验。激活单个词汇项的单令牌特征提供了可直接比较的诊断案例。我们使用全层深度的零消融分析了六个模型和三个SAE家族中的390万个特征。单令牌特征在解码器空间中聚类紧密4.7倍,并集中在早期层。消融它们在208个全层条件中的178个条件下产生了Benjamini-Hochberg显著的对数几率降低,深度控制着损害是否向下游级联或直接塑造输出。跨家族因果差异超过家族内规模效应。同一消融后,目标令牌的排名在96-98%的时间内恢复到基线的2倍以内,控制激活函数比较在同一模型内反转符号,训练方法成为剩余候选因素。因此,跨家族可解释性主张对训练方法敏感,而不仅仅是激活函数或规模。

英文摘要

Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet nobody has tested whether a feature's causal role is stable across SAE families. Single-token features fire on one vocabulary item, so ground truth permits direct comparison. We analyze 3.9M features across six models and three SAE families and zero-ablate at full layer depth: they sit 4.7x tighter in decoder space and concentrate in early layers. Deleting one lowers the model's logit for that token in 178 of 208 layer conditions, significant after multiple-comparison correction. And depth decides how the damage lands: early-layer deletions disrupt the layers that follow, late-layer deletions change the output directly. Cross-family causal differences exceed within-family scale effects: on the same base model, GemmaScope and BatchTopK features are causally anchored, LlamaScope features locally redundant. Under LlamaScope the token returns to within 2x its pre-ablation rank 96-98% of the time. Changing only the activation function reverses the sign of that difference, so the training recipe is the remaining candidate: cross-family claims are sensitive to training methodology, not just activation function or scale.

Comments23 pages, 10 figures, 27 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑