arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillEvoLean:面向 Lean 证明器的突变增强技能进化

SkillEvoLean: Mutation-enhanced skill evolution for Lean provers

Kuo Zhou, Zixiong Yang, Lu Zhang

arXiv 2610.01799首次发表:更新:

发表机构

Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出突变增强的技能自进化框架SkillEvoLean,通过渐进式与突变式更新共同进化求解策略和参考知识,在多个数学基准上显著提升Lean证明成功率。

AI 中文摘要

技能进化为在不更新大语言模型智能体参数的情况下改进其性能提供了一种有前景的途径,但其在形式化定理证明中的应用仍未得到充分探索。现有方法主要针对自然语言推理,通过分析成功和失败的轨迹并逐步修订求解策略来改进技能。尽管 Lean 验证器提供了可靠的执行反馈,但当所有采样轨迹均失败时,现有技能进化方法缺乏可用于推断有效更新方向的成功轨迹。此外,这些方法还主要关注根指令文件,从而未能充分探索包括数学概念和证明技术在内的参考知识的进化。为解决这些局限,我们提出了一种突变增强的技能自进化框架,用于构建技能增强的 Lean 证明器。该框架通过渐进式更新和基于突变的更新,共同进化高层求解策略及其参考知识。渐进式进化从成功和失败的轨迹中获取局部改进,而当无法生成完整证明时则触发突变,在验证器反馈下采样数学概念以生成并选择新的技能候选。我们在 MiniF2F、PutnamBench、2025 年国际数学奥林匹克(IMO 2025)和 2026 年美国数学奥林匹克(USAMO 2026)上评估了我们的方法。在相同的骨干模型、轨迹采样预算和测试时计算下,我们的方法在 GPT-5.5 上分别实现了 100.0%、90.6%、4/6 和 4/6 的证明成功率,优于基线方法。进一步分析表明,概念引导的突变在 MiniF2F 和 PutnamBench 上分别比随机文本引导的突变高出 6.9 和 8.2 个百分点,同时在 IMO 2025 和 USAMO 2026 上各多解决一道问题。

英文摘要

Skill evolution offers a promising way to improve large language model agents without updating their parameters, but its use in formal theorem proving remains underexplored. Existing methods mainly target natural-language reasoning, improving skills by analyzing successful and failed trajectories and incrementally revising solving strategies. Although the Lean verifier provides reliable execution feedback, when all sampled trajectories fail, existing skill evolution methods lack successful trajectories from which to infer effective update directions. Furthermore, these methods also focus mainly on the root instruction file, thus underexploring the evolution of reference knowledge including mathematical concepts and proving techniques. To address these limitations, we propose a mutation-enhanced skill self-evolution framework for building skill-augmented Lean provers. The framework jointly evolves a high-level solving policy and its reference knowledge through progressive and mutation-based updates. Progressive evolution derives local improvements from successful and failed trajectories, while mutation is triggered when no complete proof can be generated, sampling mathematical concepts to produce and select new skill candidates under verifier feedback. We evaluate our method on MiniF2F, PutnamBench, the 2025 International Mathematical Olympiad (IMO 2025), and the 2026 USA Mathematical Olympiad (USAMO 2026). Under the same backbone model, trajectorysampling budget, and test-time compute, our method achieves proof success rates of 100.0%, 90.6%, 4/6, and 4/6, respectively, with GPT-5.5, outperforming the baseline methods. Further analysis shows that concept-guided mutation outperforms random-text-guided mutation by 6.9 and 8.2 percentage points on MiniF2F and PutnamBench, respectively, while solving one additional problem on both IMO 2025 and USAMO 2026.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑