arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37991cs.AIq-bio.NC

哪些注意力头像人类大脑?不是那些进行计算的头

Which Attention Heads are like the Human Head? Not the Ones that Compute

Christopher Pinier, Gustaw Opiełka, Hannes Rosenbusch, Taylor Webb, Michael D. Nunez, Claire E. Stevenson

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在抽象模式完成任务上发现,LLM注意力头与人类大脑的对齐主要反映刺激读取方式,而非因果计算;按函数向量移除比按大脑对齐移除破坏性更大。

中文摘要 AI 辅助

大脑-人工智能对齐常被解释为模型和大脑执行相似计算的标志。但对齐单元是否因果性地参与模型计算却很少被检验。在一个抽象模式完成任务(AAABAAA → B)上,我们比较了LLM注意力头表示与人类脑电图(EEG),并测试消融这些头对任务性能的影响。对齐与因果性分离:大脑对齐的头对性能有贡献,但移除它们造成的破坏性远小于移除通过归因修补(attribution patching)选出的头。我们比较了先前可解释性工作定义的两种不参考大脑的头集:概念向量(CVs),它跨格式表示抽象模式;以及函数向量(FVs),根据其对正确答案预测的贡献进行选择。大脑对齐与FV分数几乎没有关联,而与CV分数的关联因模型而异。在大脑对齐的头中,我们发现了重复出现的注意力轮廓:一种强调独特元素(新颖头),另一种强调重复元素(重复头)。新颖家族追踪显著性,并关注与人类注视相同的元素,但其移除平均而言比随机消融破坏性更小。重复头对性能贡献适中,并与抽象模式表示(CVs)相关。在跨越3B-72B参数的17个模型中,按FV排序的移除比按大脑排序的移除破坏性显著更大。因此,大脑对齐捕捉了模型如何读取刺激,而仅微弱地捕捉了它如何表示模式并解决任务。

英文摘要

Brain-AI alignment is often interpreted as a sign that model and brain perform similar computations. Whether the aligned units are causally involved in model computation is rarely checked. On an abstract pattern-completion task (AAABAAA $\rightarrow$ B), we compare LLM attention-head representations with human EEG and test how ablating those heads affects task performance. Alignment and causation dissociate: brain-aligned heads contribute to performance, but their removal is substantially less disruptive than removal of heads selected via attribution patching. We compare two head sets that prior interpretability work defines without reference to the brain: concept vectors (CVs), which represent abstract patterns across formats, and function vectors (FVs), selected for their contribution to correct-answer prediction. Brain alignment shows little association with FV scores, while its association with CV scores varies across models. Among brain-aligned heads, we find recurring attention profiles: one emphasizes distinctive elements (novelty heads), the other repeating elements (repetition heads). The novelty family tracks salience and attends to the same elements that humans look at, yet its removal is less damaging than random ablation on average. Repetition heads contribute modestly to performance and are associated with abstract-pattern representation (CVs). Across 17 models spanning 3B-72B parameters, FV-ranked removal is substantially more disruptive than brain-ranked removal. Brain alignment thus captures how the model reads the stimulus, and only faintly captures how it represents the pattern and solves the task.

发表机构

  • University of Amsterdam(阿姆斯特丹大学)
  • Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑