arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38107cs.CLcs.AI

正确答案,无效轨迹:可验证的小学数学揭示了思维链轨迹的哪些信息

Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces

Ratish Puduppully, Pranabendu Misra, Paarth Iyer, Durgesh Kalwar, Vardhan Palod, Subbarao Kambhampati

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过可验证的小学数学基准iGSM,发现思维链轨迹的答案正确性与轨迹有效性在分布外会分离,且轨迹监督干预可削弱最小性证据,对思维链监控与解释提出警示。

中文摘要 AI 辅助

思维链轨迹被广泛视为模型如何得出答案的记录,为调试、智能体审计以及关于推理的主张提供依据。检验这种解读是困难的,因为自然语言的思考轨迹很少能被机械验证。我们在iGSM中重新审视这一问题,iGSM是一个合成的小学数学基准,旨在研究思考轨迹,并用于支持关于习得推理和规划的主张。关键在于,iGSM暴露了正确解决方案应使用的精确数量和依赖关系,使得生成的轨迹能够逐步以编程方式检查,并使我们能够测试正确答案是否可靠地伴随有效轨迹。我们首先评估仅在有效、最小轨迹上训练的模型。答案正确性和轨迹有效性在分布上几乎重合,但在分布外分离:在最难的实例中,31.6%的正确答案具有无效轨迹,其中超过一半通过了所有句法和算术检查,但未通过语义依赖检查。然后我们干预轨迹监督。非最小训练轨迹导致非最小输出,而用不同查询重新询问同一问题揭示了从原始查询继承的计算,削弱了最小性作为选择性规划证据的作用。在10%的训练轨迹句子中打乱词元,即使在分布外也保持了接近干净的准确率,尽管没有轨迹通过验证。交换的训练轨迹同样保持了较高的分布内准确率。我们讨论了这些发现对AI安全背景下思维链监控和解释的启示。

英文摘要

Chain-of-thought traces are widely read as records of how models reach their answers, informing debugging, agent auditing, and claims about reasoning. Testing this interpretation is difficult because natural-language thinking traces are rarely mechanically verifiable. We revisit it in iGSM, a synthetic grade-school mathematics benchmark designed to study thinking traces and used to support claims of learned reasoning and planning. Crucially, iGSM exposes the exact quantities and dependencies that a correct solution should use, allowing generated traces to be checked programmatically step by step and enabling us to test whether correct answers are reliably accompanied by valid traces. We first evaluate models trained exclusively on valid, minimal traces. Answer correctness and trace validity nearly coincide in distribution but decouple out of distribution: on the hardest instances, 31.6% of correct answers have invalid traces, over half of which pass all syntactic and arithmetic checks but fail semantic dependency checks. We then intervene on trace supervision. Non-minimal training traces induce non-minimal outputs, while re-asking the same problem with a different query reveals computations inherited from the original query, weakening minimality as evidence of selective planning. Shuffling tokens in 10% of training trace sentences preserves near-clean accuracy even out of distribution despite no trace passing verification. Swapped training traces likewise retain high in-distribution accuracy. We discuss the implications of these findings for chain-of-thought monitoring and interpretation in the context of AI safety.

发表机构

  • IT University of Copenhagen(哥本哈根信息技术大学)
  • Chennai Mathematical Institute(金奈数学研究所)
  • Indian Institute of Technology Jammu(印度理工学院贾姆穆分校)
  • Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑