arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24760cs.AI

构建逆向思维:发展大语言模型的逆向思考能力

Construting Reverse Thinking: Developing Large Language Models' Reverse Thingking Ability

  • China University of Petroleum (East China)(中国石油大学(华东))
  • College of Cryptology and Cyber Science, Nankai University(南开大学网络空间安全学院)
  • Goertek Inc.(歌尔股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Xin Liu, Yunhai Li, Chunfu Jia, Ziliang Chen, Jisen Song

AI总结:

针对大模型前向推理的局限,提出逆向推理模式构建方法,通过两阶段数据集与微调、细粒度奖励机制及平衡采样,显著提升数学证明等任务的推理效率与准确性。

AI中文摘要:

面对复杂问题时,人类倾向于针对不同的问题尝试各种思路。人类的思维模式在适应不同场景时表现出显著的灵活性。GPT-o1、GPT-o3和DeepSeek-R1采用长思维链模型,通过增加推理深度来解决复杂问题,这些模型默认采用前向推理模式。我们对不同规模模型在不同数学问题数据集上的准确率进行了统计分析,发现了五个错误原因:解空间覆盖不足、计算错误、未经验证的假设、忽略约束条件以及最大响应长度限制。为了解决上述问题,我们提出了一种逆向推理模式构建方法,旨在增强模型的逆向思维能力和动态适应性。首先,我们构建了一个由易到难两阶段的数学数据集,用于训练大模型并逐步提高其在不同难度水平上的推理能力。该数据集包含前向推理路径和逆向推理路径。同时,采用两阶段监督微调过程,逐步训练模型的逆向推理能力。此外,我们开发了一种细粒度的奖励机制,利用平滑的奖励信号来增强模型在推理过程中自主选择思维模式的能力,从而避免奖励黑客攻击。我们设计了一种线性衰减的平衡采样策略,以在训练过程中保持前向和逆向推理路径样本之间的平衡,使模型能够快速稳定地收敛。实验结果表明,我们的方法在数学证明等任务中显著提高了推理效率和准确性,为解决复杂问题提供了一种灵活高效的推理范式。

英文摘要:

When facing complex problems, humans tend to try various ideas for different issues. Human thinking patterns exhibit remarkable flexibility in adapting to diverse scenarios. GPT-o1, GPT-o3, and DeepSeek-R1 adopt long chain-of-thought models to address complex problems by increasing reasoning depth, which default to a forward reasoning mode. We conducted statistical analysis on the accuracy of different mathematical problem datasets on models of different scales, and found five reasons for errors: Insufficient solution-space coverage, Computational mistakes, Unverified assumptions, Ignoring constraint conditions, Maximum response length limitation. To address the above issues, we proposed a backward reasoning pattern construction method aimed at enhancing the model's reverse thinking ability and dynamic adaptability. First, we constructed an easy-hard two-stage Math dataset for training large models and gradually improving their inference ability at different difficulty levels. The dataset contains forward reasoning paths as well as backward reasoning paths. And a two-stage supervised fine-tuning process is applied to progressively train the model's backward reasoning capability. Furthermore, a fine-grained reward mechanism is developed, employing smoothed reward signals to strengthen the model's ability to autonomously select thinking modes during the reasoning process, thereby avoiding reward hacking. A linear-decay balanced sampling strategy is designed to maintain a balance between forward and backward reasoning path samples during training, enabling the model to converge quickly and stably. Experimental results show that our method significantly improves reasoning efficiency and accuracy in tasks such as mathematical proofs, offering a flexible and efficient reasoning paradigm for solving complex problems.

补充信息

↑