arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25366cs.AI

从装饰到承重:任务难度塑造思维链的因果作用

From Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought

发表机构R2M AI · 康奈尔大学 · 卡内基梅隆大学
查看机构详情
  • R2M AI
  • Cornell University(康奈尔大学)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Renee Jia, Di Mu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出续写式因果测试衡量思维链的承重度,发现任务难度主导其因果作用:简单任务中模型绕过推理,困难任务中错误传播,对CoT监督构成结构性挑战。

中文摘要 AI 辅助

思维链(CoT)监控只有在书面推理对答案产生因果约束时才有意义。我们引入了续写式因果测试,这是一种消融-补丁干预方法,它扰动一个推理步骤,截断思维链,并迫使模型从被破坏的前缀继续生成。该方法衡量思维链对最终答案的承重程度,这是一种区别于机制忠实度的行为学概念。在Gemma-2-9B-IT、Llama-3.1-8B-Instruct和DeepSeek-R1-Distill-Qwen-7B上,针对GSM8K、MMLU和BIG-Bench Hard数据集,思维链的承重度与模型相对任务难度相关:在简单任务上,模型会静默地绕过自身推理;在困难任务上,模型会遵循被破坏的步骤并传播错误。一项匹配的2x2分析表明,任务难度主导了扰动类型:从GSM8K到BBH多步算术,错误传播上升了16倍,而对28,584个续写的方差分解显示,98.8%的解释偏差归因于任务难度,而仅0.8%归因于扰动类型。针对推理的强化学习抑制了错误传播并压缩了梯度。一项四变体评判敏感性分析和盲法双标注者研究(n=500)表明,错误传播与非传播的标签对评判提示不变,且标注者间一致性完美(Cohen's kappa = 1.00)。这一梯度为基于CoT的监督和AI安全监控带来了结构性问题:在轨迹易于阅读的地方,它携带的信号很少;而在关键之处,错误在监控器干预之前就已传播。隐藏状态上的线性探针可以区分静默绕过、自我纠正和错误传播,但加性激活引导提供的因果控制有限,最多仅能翻转约25%的错误传播案例。行为模式可读但不可靠地可控。

英文摘要

Chain-of-thought (CoT) monitoring is only meaningful if written reasoning causally constrains the answer. We introduce continuation-based causal testing, an ablation-patch intervention that perturbs one reasoning step, truncates the chain, and forces the model to continue from the corrupted prefix. It measures how load-bearing a CoT is for the final answer, a behavioral notion distinct from mechanistic faithfulness. Across Gemma-2-9B-IT, Llama-3.1-8B-Instruct, and DeepSeek-R1-Distill-Qwen-7B on GSM8K, MMLU, and BIG-Bench Hard, CoT load-bearingness tracks model-relative task difficulty: on easy tasks models silently bypass their own reasoning; on hard tasks they follow corrupted steps and propagate errors. A matched 2x2 analysis shows task difficulty dominates perturbation type: error propagation rises 16x from GSM8K to BBH multistep arithmetic, and a variance partition over 28,584 continuations attributes 98.8% of explained deviance to task difficulty versus 0.8% to perturbation type. Reasoning-specific RL suppresses error propagation and compresses the gradient. A four-variant judge-sensitivity analysis and blind two-annotator study (n=500) show the error-propagation vs. non-propagation label is invariant to judge prompt, with perfect inter-annotator agreement (Cohen's kappa = 1.00). This gradient creates a structural problem for CoT-based oversight and AI safety monitoring: where the trace is easy to read it carries little signal, and where it matters errors propagate before a monitor can intervene. Linear probes on hidden states separate silent bypass, self-correction, and error propagation, but additive activation steering provides limited causal control, flipping only about 25% of error-propagation cases at best. Behavioral mode is readable but not reliably controllable.

补充信息

↑