发表机构
İzmir Institute of Technology(伊兹密尔理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出顺序激活补丁框架及顺序多头补丁,探究思维链推理的因果效应位置与相关注意力头,证实存在支持思维链的分布式推理子电路。
AI 中文摘要
大型语言模型(LLMs)在思维链(Chain-of-Thought, CoT)提示引导下展现出卓越的问题解决能力,但这类改进背后的内部机制仍鲜为人知。本研究探究与CoT相关的因果效应在生成的推理轨迹中何处出现,以及哪些注意力头携带有助于最终答案计算的信号。由于CoT推理在多个生成的token上展开,标准的单静态token位置激活补丁不足以表征这些时间分布的效应。为解决这一局限,我们引入顺序激活补丁框架,该框架追踪token位置上CoT条件注意力头的激活,并使用词性引导分析聚合其效应;我们还引入顺序多头补丁,以评估分布式头集的联合贡献,同时结合跨问题和随机激活控制。针对性的零消融实验表明,识别出的头对成功生成答案具有功能重要性,并影响多种重叠机制,包括推理轨迹维持、答案锚定、示例-目标分离及数值生成。总体而言,我们的结果为与CoT条件计算相关的分布式推理支持子电路提供了证据。
英文摘要
Large Language Models (LLMs) demonstrate remarkable problem-solving capabilities when guided by Chain-of-Thought (CoT) prompting, yet the internal mechanisms underlying these improvements remain poorly understood. In this work, we investigate where CoT-related causal effects emerge across the generated reasoning trajectory and which attention heads carry signals that contribute to final-answer computation. Because CoT reasoning unfolds over multiple generated tokens, standard activation patching at a single static token position is insufficient to characterize these temporally distributed effects. To address this limitation, we introduce a sequential activation patching framework that traces CoT-conditioned attention-head activations across token positions and aggregates their effects using Part-of-Speech-guided analysis. We further introduce Sequential Multi-Head Patching to evaluate the joint contribution of distributed head sets, together with cross-question and random activation controls. Targeted zero-ablation experiments show that the identified heads are functionally important for successful answer generation and affect several overlapping mechanisms, including reasoning-trajectory maintenance, answer anchoring, exemplar-target separation, and numerical generation. Overall, our results provide evidence for distributed reasoning-support sub-circuits associated with CoT-conditioned computation.