发表机构
University of Manchester; Idiap Research Institute(曼彻斯特大学; Idiap研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出融贯论概率组合主义(CPC)框架,将Transformer计算分为对齐、统一、抑制、路由四种算子角色,在15个模型上验证其有效性,为Transformer机制比较提供通用术语。
AI 中文摘要
机械可解释性已识别出Transformer电路,但缺乏用于描述其在不同任务和架构间功能组合的通用术语。我们提出融贯论概率组合主义(Coherentist Probabilistic Compositionalism, CPC),这一解释框架将Transformer计算建立在融贯论解释理论基础上,通过四种算子角色描述其计算过程:对齐(Alignment)识别候选关系,统一(Unification)整合支撑信息,抑制(Suppression)减少不相容备选,路由(Routing)将选定信息传递至输出。在来自5个架构家族的15个模型中,抑制、统一和路由的权重空间特征与保留的激活层面角色度量的相关性高于随机基线;抑制在不同任务间的稳定性优于统一。在10个模型中,对对齐头(alignment heads)进行消融会降低下游抑制活性,超出随机头对照的效果;但对无冲突提示的类似影响表明,这是一种普遍的上游依赖,而非矛盾特定的耦合。显式矛盾会显著改变14个模型的分层一致性代理;在移除共享残差协方差后,所有模型的差距均符合预期方向。基础模型和指令调优变体保留了归纳头得分结构(r≥0.98),且算子特征未向更深层出现一致偏移。这些结果支持CPC作为比较Transformer机制的通用术语,同时表明其深度和几何表达仍具有架构特异性。
英文摘要
Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures. We introduce Coherentist Probabilistic Compositionalism (CPC), an interpretive framework that grounds transformer computation in coherentist theories of interpretation and describes it through four operator roles. Alignment identifies candidate relations, unification integrates supporting information, suppression reduces incompatible alternatives, and routing carries selected information to the output. Across 15 models from five architecture families, the suppression, unification, and routing weight-space signatures correlate with held-out activation-level role measures above random baselines. Suppression is more stable across tasks than unification. Ablating alignment heads reduces downstream suppressive activity beyond a random-head control in 10 models, but similar effects on no-conflict prompts indicate a general upstream dependency, not contradiction-specific coupling. Explicit contradictions significantly shift a layerwise coherence proxy in 14 models; after removing shared residual covariance, the gap has the predicted direction in every model. Base and instruction-tuned variants preserve induction-head score structure ($r{\geq}0.98$) without a consistent shift of operator signatures towards later layers. These results support CPC as a shared vocabulary for comparing transformer mechanisms while showing that their depth and geometric expression remain architecture-specific.