arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00373cs.LGcs.CL

注意力头消融何时支持因果主张?投影层混杂、地板效应与匹配对照

When Do Attention-Head Ablations Support Causal Claims? Projection-Level Confounds, Floor Effects, and Matched Controls

Juli Huang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过GPT-2 small和DistilGPT2实验表明,注意力头消融的因果推断需注意干预位置、连续指标和匹配对照,否则结论可能不可靠。

中文摘要 AI 辅助

注意力头消融,即将某个头置零并测量由此导致的任务性能变化,是推断语言模型中哪些组件对某一行为具有因果责任的常用方法。我们使用GPT-2 small证明,除非干预语义、评估指标和对照设置经过仔细验证,否则这种推断可能很脆弱。一种自然的投影后“置零头”实现与修正后的投影前消融几乎不相关(皮尔逊r = 0.057),并且选择出完全不相交的前5个重要头集合。我们还表明,二元准确率可能掩盖行为地板和接近天花板处的效应,而金标词元对数概率则保持梯度性。使用发现/留出划分和1,000次匹配的随机头与层匹配头对照抽样,修正后的逐头效应排名在不同划分间高度稳定(斯皮尔曼rho = 0.974),且前5个选中的头显著超过两种对照分布(蒙特卡洛p = 0.001)。然而,任务特异性的证据在GPT-2上并不稳健。在DistilGPT2上的复现保留了干预语义和匹配对照的发现。这些结果表明,单头消融本身并不能证明因果主张;可辩护的解释需要正确的干预位置、非饱和的连续指标以及匹配的留出对照。

英文摘要

Attention-head ablation, zeroing a head and measuring the resulting change in task performance, is a common method for inferring which components of a language model are causally responsible for a behavior. We show using GPT-2 small that this inference can be fragile unless the intervention semantics, evaluation metric, and controls are carefully validated. A natural post-projection implementation of "zeroing a head" is nearly uncorrelated with a corrected pre-projection ablation (Pearson r = 0.057) and selects a completely disjoint top-5 set of important heads. We also show that binary accuracy can hide effects at behavioral floors and near ceilings, whereas gold-token log-probability remains graded. Using a discovery/held-out split and 1,000 matched random-head and layer-matched-head control draws, the corrected per-head effect ranking is highly stable across splits (Spearman rho = 0.974), and the top-5 selected heads significantly exceed both control distributions (Monte Carlo p = 0.001). However, evidence for task specificity is not robust on GPT-2. Replication on DistilGPT2 preserves the intervention-semantic and matched-control findings. These results show that single-head ablation does not by itself justify a causal claim; defensible interpretation requires correct intervention placement, a non-saturated continuous metric, and matched held-out controls.

发表机构

  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑