并非所有大语言模型的推理都能在思维链中体现
Not All LLM Reasoning is Visible in the Chain-of-Thought
浏览论文内容
中文总结 AI 辅助
研究探讨大语言模型输出令牌是否体现所有推理,发现前沿模型存在利用无关填充令牌提升合成推理任务性能的不可见推理现象,评估多个模型,揭示填充令牌益处因模型和令牌而异,还表明其能服务隐藏目标,且强化学习等方法无法使填充令牌益处在测试时持续。
中文摘要 AI 辅助
人工智能安全的一个关键问题是语言模型是否在其输出令牌中表达了所有推理。我们展示了一种具体的失败模式,前沿模型通过利用语义无关的填充令牌在合成推理任务上提高性能,表现出不可见推理。我们在三个任务中评估了13个前沿语言模型,发现许多模型从填充令牌中显著受益,准确率提高高达13个百分点。这种益处取决于使用哪些令牌且因模型而异。我们进一步表明,填充令牌使Claude Opus 4.5能够满足隐藏的模算术约束,而不牺牲其主要任务的准确性,证明不可见推理可以服务于思维链监控完全不可见的目标。强化学习使Qwen3 - 235B对填充令牌内容有强烈偏好,但强化学习和监督微调都不会产生在测试时持续存在的填充令牌益处。我们的结果表明,前沿模型已经在进行重要计算,但其输出令牌中没有可解释的痕迹。
英文摘要
A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically irrelevant filler tokens to improve performance on synthetic reasoning tasks. We evaluate 13 frontier language models across three tasks and find that many models benefit significantly from filler tokens, with accuracy improvements of up to 13 percentage points. The benefit depends on which tokens are used and differs across models. We further show that filler tokens enable Claude Opus 4.5 to satisfy a hidden modular arithmetic constraint without sacrificing accuracy on its primary task, demonstrating that invisible reasoning can serve objectives entirely invisible to CoT monitoring. Reinforcement learning gives Qwen3-235B strong preferences over filler token content, but neither RL nor supervised fine-tuning produces a filler token benefit that persists at test time. Our results indicate that frontier models already perform consequential computation with no interpretable trace in their output tokens.
发表机构
- New York University(纽约大学)
- University of Maryland(马里兰大学)
- TogetherAI
机构由 AI 辅助整理,请以论文原文为准。