arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28600cs.AI

数学推理中思维链的SHAPE

SHAPE of Chain-of-Thought in Math Reasoning

Jonghyun Song, Sangjun Song, Minjae Oh, Haesung Pyun, Sungsik Lee, Yohan Jo

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出用于分析LLM思维链轨迹的\texttt{SHAPE}框架,发现数学启发式比传统CoT特征更能解释答案正确性,基于此的后训练可提升LLM数学推理准确率。

中文摘要 AI 辅助

大型语言模型(LLM)在数学推理基准测试中表现出色,但支撑其推理的具有数学意义的技能仍未得到充分探索。我们引入\texttt{SHAPE},这是一个通过数学教育领域开发的两个视角分析思维链(CoT)轨迹的框架:(1)语义空间:模型对问题不断演变的数学解释(例如代数、几何);(2)启发式:在这些空间内采取的具体数学行动(例如简化问题、反向推导)。我们首先使用\texttt{SHAPE}分析各种模型的推理模式,发现模型采用的数学启发式比传统的CoT特征更能解释最终答案的正确性;此外,模型更可能通过将推理精力集中在少数语义空间而非探索多个不同空间来获得正确解决方案,这一模式与人类行为一致。接下来,我们利用\texttt{SHAPE}视角评估后训练是否真正提升了数学能力,发现强化学习会导致启发式使用的模式寻求行为。最后,我们通过促进多样化的启发式对LLM进行后训练,证明其在提高准确率方面的有效性。总体而言,\texttt{SHAPE}为解码LLM推理提供了一个基于理论的诊断框架,并为后训练LLM以提升数学推理开辟了新路径。模型代码可在该httpsURL获取。

英文摘要

Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically meaningful skills underlying their reasoning remain underexplored. We introduce \texttt{SHAPE}, a framework that analyzes Chain-of-Thought (CoT) trajectories through two lenses developed in mathematics education: (1) semantic spaces: the model's evolving mathematical interpretations of a problem (e.g., algebraic, geometric), and (2) heuristics: the specific mathematical actions taken within those spaces (e.g., simplifying the problem, working backward). We first use \texttt{SHAPE} to analyze the reasoning patterns of various models. Our findings reveal that the mathematical heuristics employed by a model better explain final answer correctness than traditional CoT features. Furthermore, models are likely to reach correct solutions by concentrating their reasoning effort within a few semantic spaces rather than exploring many disparate ones -- a pattern consistent with human behavior. Next, we utilize the \texttt{SHAPE} lens to evaluate whether post-training truly enhances mathematical proficiency. We find that reinforcement learning induces mode-seeking in heuristic usage. Lastly, we post-train LLMs by promoting diverse heuristics and demonstrate its effectiveness in improving accuracy. Overall, \texttt{SHAPE} provides a theoretically-grounded diagnostic framework for decoding LLM reasoning and offers a new path toward post-training LLMs for math reasoning. The code for our model is available at https://github.com/holi-lab/SHAPE-of-CoT

补充信息

↑