发表机构
Bundesdruckerei GmbH(联邦印刷公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出事后分层结构,通过自我改进循环联合训练推理模型的三种能力,给出方法形式化规范及Lean定理证明的实例,实证评估待开展。
AI 中文摘要
我们基于以下观察为推理模型引入了一种自我改进循环:即使问题难度超出模型当前的求解能力,额外提供的解决方案仍可能让模型事后提取有用的解题思路。我们通过联合训练同一模型以展现以下三种能力来实现这一点:仅从问题预测解题思路、从问题和已知方案逆向推导思路、利用提供的思路求解问题。该循环交替进行:从带有提供方案的问题中逆向推导思路,并将这些思路作为额外监督信号,用于联合训练所有三种能力。我们给出了该方法的形式化规范及在Lean定理证明器中用于交互式定理证明的具体实例;实证评估仍为未来工作。
英文摘要
We introduce a self-improvement loop for reasoning models based on the following observation: Even when the difficulty of a problem exceeds the model's current solving abilities, an additionally supplied solution might enable the model to extract useful solution ideas in hindsight. We operationalize this by jointly training the same model to exhibit the following three capabilities: predicting solution ideas from problems alone, reverse-engineering ideas from problems and known solutions, and solving problems using provided ideas. The loop alternates between reverse engineering such ideas from problems with supplied solutions and using these ideas as additional supervision for joint training of all three capabilities. We give a formal specification of our method and a concrete instantiation for interactive theorem proving in the Lean theorem prover; empirical evaluation remains future work.
Comments21 Pages, 4 Figures