arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

INSPIRE:一种用于示例驱动数学推理的“先内化再改进”方法

INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning

Shuai Wang, Jiayi Kuang, Yinghui Li, Haojing Huang, Xinnian Liang, Ying Shen, Liang Lin

arXiv 2608.27501首次发表:更新:

发表机构

Sun Yat-sen University; Tsinghua University; ByteDance Inc.; Peng Cheng Laboratory(中山大学; 清华大学; 字节跳动公司; 鹏城实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLMs数学推理中内化不足的问题,提出INSPIRE方法,结合RGSI与分阶段规则偏好训练,提升了示例驱动数学推理能力且未降低通用推理水平。

AI 中文摘要

大型语言模型(LLMs)在数学推理方面取得了快速进展,但现有方法主要针对最终答案的正确性进行优化,这引发了一个问题:模型是真正内化了数学概念,还是仅仅记住了解题模式。在人类数学教育中,基于示例的推理(如构造反例以测试定理边界)反映了深刻的概念理解,但在当前的LLMs中仍未得到充分发展。通过偏好优化来增强这种能力面临两个关键挑战:(1)模型有限的基于示例的推理能力使得构建有效的偏好对本质上很困难;(2)能力获取是渐进式的,模型必须先学会采用这种策略,然后才能学会正确应用它。因此,我们提出了INSPIRE,一种“先内化再改进”的方法,结合了参考引导的学生内化(RGSI,在策略模型自身分布下生成高质量偏好候选)和分阶段的规则偏好训练策略(将学习分解为面向方法和面向正确性的阶段)。在多个模型规模和系列上的实验显示出一致的改进,甚至超过了更大的开源模型,而对分布外基准的评估证实,模型的一般数学推理能力没有下降。

英文摘要

Mathematical reasoning has seen rapid progress in large language models (LLMs), yet existing methods optimize predominantly for final-answer correctness, raising the question whether models truly internalize mathematical concepts or merely memorize solution patterns. In human mathematics education, example-based reasoning such as constructing counterexamples to test theorem boundaries reflects deep conceptual understanding, but remains underdeveloped in current LLMs. Enhancing this capability through preference optimization presents two key challenges: (1) the model's limited example-based reasoning ability makes constructing effective preference pairs inherently difficult; and (2) capability acquisition is progressive, as the model must first learn to adopt this strategy before learning to apply it correctly. Therefore we propose INSPIRE, an Internalize-Then-Improve approach combining Reference-Guided Student Internalization (RGSI), which produces high-quality preference candidates under the policy model's own distribution, with a stage-wise rubric preference training strategy that decomposes learning into method-oriented and correctness-oriented stages. Experiments across multiple model scales and families demonstrate consistent improvements, even surpassing larger open-source models, while evaluations on out-of-distribution benchmarks confirm no degradation in general mathematical reasoning ability.

CommentsEMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑