CASE:用于提高思维链忠实度的因果对齐和结构强化
CASE: Causal Alignment and Structural Enforcement for Improving Chain-of-Thought Faithfulness
浏览论文内容
中文总结 AI 辅助
研究思维链推理中生成的推理无法忠实支持最终答案的问题,提出CASE框架,通过训练时因果对齐和推理时结构强化,结合构建多种数据集及选择性损失微调、屏蔽注意力等方法,提升CoT忠实度、跨数据集转移能力及平均准确率。
中文摘要 AI 辅助
思维链(CoT)推理被广泛用于提升大语言模型(LLMs)的性能和可解释性,但生成的推理可能无法忠实支持最终答案。本文从因果角度研究该问题,忠实的CoT过程应遵循\(Z\rightarrow X\rightarrow Y\)的链条,其中\(Z\)、\(X\)、\(Y\)分别表示指令、推理链和最终答案。传统自回归LLMs在指令和CoT上进行答案生成,存在直接的指令到答案的捷径。为此提出CASE框架,在训练时构建反事实CoT、有偏指令和空指令数据集,并应用选择性损失微调来加强CoT到答案的依赖,同时抑制指令捷径。在推理时,CASE屏蔽从指令令牌到答案令牌的直接注意力,防止模型绕过生成的CoT。通过信息论分析展示了这些组件如何促进忠实链条。在三个模型和四个基准上的实验表明,CASE在整体CoT忠实度上比最强基线平均提高了37%,具有更强的跨数据集忠实度转移,并保持了有竞争力的平均准确率。
英文摘要
Chain-of-thought (CoT) reasoning is widely used to improve both the performance and interpretability of large language models (LLMs), yet the generated reasoning may not faithfully support the final answer. We study this problem from a causal perspective, where a faithful CoT process should follow the chain $Z\rightarrow X\rightarrow Y$, with $Z$, $X$, and $Y$ denoting the instruction, reasoning chain, and final answer, respectively. In this process, the instruction should affect the answer only through the reasoning chain. However, conventional autoregressive LLMs condition answer generation on both the instruction and the CoT, which still allows a direct instruction-to-answer shortcut. To address this issue, we propose CASE, a framework that combines training-time causal alignment and inference-time structural enforcement. During training, CASE builds counterfactual-CoT, biased-instruction, and empty-instruction datasets, and applies selective-loss fine-tuning to strengthen CoT-to-answer dependence while suppressing instruction shortcuts. During inference, CASE masks direct attention from instruction tokens to answer tokens, preventing the model from bypassing the generated CoT. We provide an information-theoretic analysis showing how these components promote faithful chains. Experiments on three models and four benchmarks show that CASE achieves a 37\% average per-setting relative improvement in overall CoT faithfulness over the strongest baselines, exhibits stronger cross-dataset faithfulness transfer, and maintains competitive average accuracy. Code is available at https://github.com/oddwang/CASE.