arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

思维链不忠实的两种模式:模型错误时行为检测失效

Two Regimes of Chain-of-Thought Unfaithfulness: Metric-Based Detection Fails Where Models Are Wrong

Suramya R. Angdembay, Dikshant Aryal, Nick Rahimi

arXiv 2607.23458首次发表:更新:

发表机构

The University of Southern Mississippi(南密西西比大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究思维链不忠实检测,发现答案正确性影响检测效果,分正确与错误答案两种模式。仅答案不正确性表现优,跨模式无共享正向对齐方向,指示轨迹难转移,还解决了基准标签语义不匹配问题。

AI 中文摘要

思维链(CoT)解释只有在忠实的情况下才能支持监督,即所述推理必须实际得出答案。通过针对FaithCoT-Bench的人工注释审计不忠实CoT的黑盒(行为)检测,我们发现答案正确性在各个层面构建了问题。仅答案不正确性(一种预言诊断,而非可部署的检测器)就优于所有专门构建的信号(AUROC 0.696),因为69%的注释不忠实发生在错误答案上。按正确性分层将检测分为两种模式:在正确答案上,行为信号能适度区分忠实推理和事后推理(0.63 - 0.67);在错误答案上,大多数不忠实情况存在于此,没有测试信号能明显高于随机水平(在所有四个模型的基准范围内信号上得到复制)。标准的步骤去除度量与人工标签呈反相关;这种反转在基准发布的分数和依赖提示的反事实标记轨迹上重现。线性探针解码了Llama - 3.1 - 8B中行为盲视模式和Qwen - 2.5 - 7B中正确答案模式,跨模式未检测到共享的正向对齐方向;指示的答案优先轨迹(7个模型)在两种注释模式下均无法转移,而提示诱导的未言语化答案翻转在依赖模型和源的设置中可以转移。我们还独立验证并解决了基准标签语义中的文档 - 数据不匹配问题。

英文摘要

Chain-of-thought (CoT) explanations support oversight only if they are faithful: the stated reasoning must actually produce the answer. Auditing black-box (behavioral) detection of unfaithful CoT against FaithCoT-Bench's human annotations, we find answer correctness structures the problem at every level. Answer incorrectness alone (an oracle diagnostic, not a deployable detector) outperforms every purpose-built signal (AUROC 0.696), because 69% of annotated unfaithfulness occurs on incorrect answers. Stratifying by correctness splits detection into two regimes: on correct answers, behavioral signals moderately separate faithful from post-hoc reasoning (0.63-0.67); on incorrect answers, where most unfaithfulness lives, no tested signal is detectably above chance (replicated on all four models for benchmark-wide signals). The standard step-removal metric anti-correlates with human labels; this inversion reproduces on the benchmark's released scores and on hint-dependent counterfactually labeled traces. Linear probes decode the behaviorally blind regime in Llama-3.1-8B and the correct-answer regime in Qwen-2.5-7B, with no shared, positively aligned direction detected across regimes; instructed answer-first traces (7 models) transfer to neither annotated regime, while hint-induced unverbalized answer flips do, in model- and source-dependent settings. We also independently verify and resolve a documentation-data mismatch in the benchmark's label semantics.

Comments14 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑