AI 中文总结
本研究通过机制分析识别出语言模型中的承诺-弃权回路(CAC),揭示其过度承诺导致幻觉的机制,并利用该回路训练轻量策略,显著提升决策准确率并减少虚假弃权。
AI 中文摘要
语言模型(LMs)常常通过给出自信的回答而非弃权(不执行)来产生幻觉,即使它们没有足够的信息来可靠地回答。现有的大量工作通过检测或弃权机制来缓解幻觉,但未阐明模型在内部如何首先做出承诺或弃权的决定。我们通过机制分析研究这一决定,将幻觉视为无依据的承诺:模型尽管表现出不可回答性的信号,仍然做出承诺。利用因果门控,我们识别出一个承诺-弃权回路(CAC),这是一个稀疏、因果定位的注意力头和MLP子层子集,支撑这一决定。在来自五个家族的十个语言模型(3B-14B)和三个基准上,CAC表现出一种反复出现的“累积但纠正不足”模式:承诺促进组件在较早层累积承诺,而弃权促进组件在较后层作为纠正信号,但往往不足以推翻累积的承诺。基于这一发现,一个在CAC激活上训练的轻量级策略将决策准确率比模型固有的承诺-弃权边际提高了12.2个百分点,将虚假弃权减少了2.5倍,可迁移到未见过的基准,并扩展到更大的模型(27B-35B)。CAC既具有诊断性,阐明了模型如何过度承诺,又具有实用性,实现了改进的弃权决策。
英文摘要
Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to answer reliably. A large body of existing work mitigates hallucination through detection or abstention mechanisms, but leaves open how models internally arrive at the decision to commit or abstain in the first place. We study this decision through mechanistic analysis, framing hallucination as unsupported commitment: the model commits despite exhibiting signals of unanswerability. Using causal gating, we identify a Commit-Abstain Circuit (CAC), a sparse, causally localised subset of attention heads and MLP sublayers underlying this decision. Across ten LMs (3B-14B) from five families and three benchmarks, the CAC exhibits a recurring accumulate-yet-undercorrect pattern: commitment-promoting components build up commitment in earlier layers, while abstention-promoting components act later as corrective signals that are often insufficient to overturn the accumulated commitment. Building on this finding, a lightweight policy trained on CAC activations improves decision accuracy by 12.2 points over the model's intrinsic commit-abstain margin, reduces false abstentions by 2.5 times, transfers to unseen benchmarks, and extends to larger models (27B-35B). The CAC is both diagnostic, clarifying how models overcommit, and practical, enabling improved abstention decisions.
CommentsAccepted to NeuRIPS 2026 Main Conference