arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Matryoshka 归因:学习将语言模型输出归因于表示和权重

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

Aryaman Arora, Kirill Acharya, Nathan Hu, Yanzhe Zhang, Noah Goodman, Dan Jurafsky, Christopher Potts

arXiv 2609.25518首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出 Matryoshka 归因(MAttr)掩码学习方法,通过可微 sigmoid top-k 算子学习组件嵌套子集,在机制可解释性基准上排名第一,并可用于识别微调中导致行为变化的权重。

AI 中文摘要

将语言模型输出归因于其内部计算是可解释性中的一个开放问题。现有方法使用因果干预、梯度或可学习掩码,要么代价高昂得不可行,要么难以识别实际因果重要的内部计算。我们提出将归因问题表述为识别使下游损失最小化的内部组件嵌套子集的问题。为了学习这一任务,我们引入了 Matryoshka 归因(MAttr),一种掩码学习方法,使用简单的可微 sigmoid top-$k$ 算子对掩码进行参数化。我们通过在训练过程中随机化 $k$,同时监督所有稀疏度下的训练,从而学习到按归因分数排序的组件顺序。MAttr 在机制可解释性基准(Mueller 等人,2025)的官方排行榜上排名第一;我们的方法能够识别跨不同电路基础的稀疏且可任务迁移的电路。作为实际应用,我们展示了 MAttr 可以通过强化学习进行训练,以识别导致 LLM 微调中下游行为的权重变化。我们在拒绝评判分数上训练 MAttr,发现将 Llama 3.1 8B Instruct 权重的 $1\%$ 恢复到其基础模型状态,足以消除拒绝行为,同时保持能力。我们将 MAttr 视为将可解释性成功表述为一个可通过梯度下降解决的可学习目标,并鼓励沿着这些方向开展未来工作。

英文摘要

Attributing language model outputs to their internal computations is an open problem in interpretability. Existing methods, which use causal interventions, gradients, or learnable masks, either are infeasibly expensive or struggle to identify actual causally-important internal computations. We propose framing attribution as the problem of identifying nested subsets of internal components which minimise a downstream loss. To learn this task, we introduce Matryoshka Attribution (MAttr), a mask learning method that parametrises the mask with a simple differentiable sigmoid top-$k$ operator. We supervise training over all sparsities simultaneously by randomising $k$ over training, resulting in a learned ordering of components by attribution score. MAttr achieves number 1 on the official leaderboard of the Mechanistic Interpretability Benchmark (Mueller et al., 2025); our method identifies sparse and task-transferrable circuits across varying circuit bases. As a practical application, we show that MAttr can be trained with reinforcement learning to identify weight changes responsible for downstream behaviours in LLM finetuning. We train MAttr on refusal judge scores and find that restoring $1\%$ of Llama 3.1 8B Instruct's weights to their base model state is sufficient to remove refusals while maintaining capabilities. We view MAttr as a successful formulation of interpretability into a learnable objective that we can tackle with gradient descent, and encourage future work along these lines.

Comments10 pages main text, 58 pages total; preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑