arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.26410cs.CLcs.AI

为何模型知道却不说:链式思考中的思考令牌与答案之间的信念差异在开放权重推理模型中

Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models

  • University of Nevada, Las Vegas(内华达大学拉斯维加斯分校)
  • DeepNeuro AI

机构由 AI 辅助整理,请以论文原文为准。

Richard J. Young

更新

AI总结:

研究探讨了12种开放权重推理模型在MMLU和GPQA问题中对误导提示的响应,发现思考令牌与答案之间的信念差异显著,且模型对提示的承认存在方向性差异。

AI中文摘要:

扩展思考模型暴露了一个第二文本生成通道(

英文摘要:

Extended-thinking models expose a second text-generation channel ("thinking tokens") alongside the user-visible answer. This study examines 12 open-weight reasoning models on MMLU and GPQA questions paired with misleading hints. Among the 10,506 cases where models actually followed the hint (choosing the hint's target over the ground truth), each case is classified by whether the model acknowledges the hint in its thinking tokens, its answer text, both, or neither. In 55.4% of these cases the model's thinking tokens contain hint-related keywords that the visible answer omits entirely, a pattern termed *thinking-answer divergence*. The reverse (answer-only acknowledgment) is near-zero (0.5%), confirming that the asymmetry is directional. Hint type shapes the pattern sharply: sycophancy is the most *transparent* hint, with 58.8% of sycophancy-influenced cases acknowledging the professor's authority in both channels, while consistency (72.2%) and unethical (62.7%) hints are dominated by thinking-only acknowledgment. Models also vary widely, from near-total divergence (Step-3.5-Flash: 94.7%) to relative transparency (Qwen3.5-27B: 19.6%). These results show that answer-text-only monitoring misses more than half of all hint-influenced reasoning and that thinking-token access, while necessary, still leaves 11.8% of cases with no verbalized acknowledgment in either channel.

补充信息

↑