重连还是门控?指令微调如何塑造大语言模型中的知识冲突电路
Rewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs
浏览论文内容
中文总结 AI 辅助
本研究首次对比基础与指令微调模型的知识冲突电路,发现微调通过门控重加权而非重连电路,增强对简短反事实的拒绝,但该鲁棒性受框架限制。
中文摘要 AI 辅助
在语言模型中,相信提示词还是相信权重参数的选择由少数可识别的注意力头做出。指令微调改变了模型在冲突情境下的行为,但它究竟是重连了底层电路,还是仅仅对已存在的组件进行门控或重新加权,目前仍不清楚。我们首次对冲突解决电路进行了机制层面的基础模型与指令微调模型对比研究,涵盖三个模型家族(Llama-3.2-3B、Qwen-2.5-3B、Gemma-3-4B)。五种独立方法——节点和边归因、叠加角色分析、因果消融以及路径修补——均收敛于门控机制,且发现相同的注意力头位于相同的深层,被重新加权而非替换,节点重叠度较高(0.60-0.82)。在行为层面,微调使模型偏向参数记忆,使得指令微调模型拒绝简短的反事实上下文的程度远高于基础模型,这与天真的用户跟随预期相反。然而,这种增加的怀疑态度是框架因素所致,因为当相同的虚假声明以连贯、有证据的段落形式呈现时,这种怀疑态度消失。因此,指令微调带来的针对简短注入的鲁棒性是真实但狭窄的。更广泛地说,我们认为由于冲突电路被保留而非重建,基于基础模型校准的可解释性和控制工具应能直接迁移到其部署的指令微调版本上。
英文摘要
In language models, the choice between believing the prompt and believing the weights is made by a handful of identifiable attention heads. Instruction tuning changes how models behave under conflict, but whether it rewires the underlying circuit or merely gates/reweights already present components, remains unknown. We provide the first mechanistic base-vs-instruct comparison of conflict-resolution circuits, across three families (Llama-3.2-3B, Qwen-2.5-3B, Gemma-3-4B). Five independent methods, node and edge attribution, superposition role analysis, causal ablation, and path patching, converge on gating, with the same heads, in the same late-layers, are found to be reweighted rather than replaced with a high node overlap (0.60-0.82). Behaviorally, tuning shifts models toward parametric memory, making instruct models reject a terse counterfactual context far more than base ones, the opposite of a naive user-following expectation. Yet this added skepticism is a factor of framing since it disappears when the same false claim is delivered as a coherent, evidential passage. The robustness that instruction tuning buys against terse injection is therefore real but narrow. More broadly, we believe that because the conflict circuit is preserved rather than rebuilt, interpretability and control tools calibrated on base models should transfer directly to their deployed instruct siblings.
发表机构
- IvLabs, VNIT(伊夫实验室,维萨拉亚国家理工学院)
机构由 AI 辅助整理,请以论文原文为准。