arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22007cs.LGcs.CL

Transformer的通信图

The Communication Map of a Transformer

  • St. John Fisher University(圣约翰费舍尔大学)

机构由 AI 辅助整理,请以论文原文为准。

Richard Zhe Wang

AI总结:

该研究提出仅从权重生成语言模型所有潜在通信通道的通信图,通过实验验证其可恢复诱导电路并消除模型的诱导能力,发布了相关工具套件。

AI中文摘要:

Transformer的组件通过向共享残差流写入和读取进行通信,机制可解释性已手动绘制了这些连接,每次一个电路。我们提出了通信图,它仅从权重出发绘制语言模型中的所有潜在通信通道,将Elhage等人(2021)的组合分数推广为涵盖全部18种连接类别的单个耦合系数,范围从整个注意力头电路到单个神经元。对所有候选通道的统计显示,GPT-2中有6.3×10^8个,Pythia-6.9B中有1.3×10^11个,发现70%-89%的头对的取向远离随机,一些耦合紧密,另一些则主动相互回避。完整图在单个消费级GPU上,GPT-2需15秒,Pythia-6.9B需11分钟。两个应用展示了该图的实用性:应用1中,最强的头对头耦合能在无先验信息的情况下恢复已知的诱导电路,并将其分组为社区,消融其中一个社区会破坏模型的上下文复制能力;应用2中,汇集每个头的耦合系数可识别出一个独特的二维流子空间,删除该子空间会消除包括Pythia-6.9B在内的六个模型的诱导能力,该子空间与激活主成分分析或异常值维度识别的子空间不同。我们发布了该图、统计机制和干预套件。

英文摘要:

The components of a transformer communicate by writing to and reading from a shared residual stream, and the mechanistic interpretability literature has mapped these connections by hand, one circuit at a time. We present the communication map, which charts every potential communication channel from the geometry of the model's weights alone, generalizing the composition score of Elhage et al. (2021) into a single coupling coefficient covering all 18 connection classes, from head-to-head to neuron-to-neuron and everything in between. We provide an account of the properties of the coupling coefficient, including its geometric interpretation and its exact chance level. The census finds that 70-89% of head pairs are oriented far from chance, some coupled strongly and others actively avoiding each other. We demonstrate the communication map in two novel applications. In Application 1, we recover the known induction circuits blind from the strongest head-to-head couplings and group the heads into communities, and ablating one such community destroys the model's in-context copying. In Application 2, we pool the coupling coefficients of every head to identify a distinct two-dimensional residual stream subspace, whose deletion abolishes the induction capability in six models up to Pythia-6.9B. We show that this subspace is different from those identified by either activation PCA or outlier dimensions. We release the map, the statistical machinery, and the intervention suite.

补充信息

↑