arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于验证器的基准协同进化的自修改精益证明智能体

Self-Modifying Lean Proof Agents with Verifier-Grounded Benchmark Coevolution

Yuqing Li, Zeguan Wu, Yu Gan, Junyu Liu

arXiv 2607.17352首次发表:更新:

发表机构

University of Pittsburgh(匹兹堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究设计有效精益证明智能体的挑战,提出自进化精益证明智能体,其工作区可变,与基准协同进化,通过特定更新机制保持分数可比,经实验验证该方法能有效改善精益证明工作流程。

AI 中文摘要

设计有效的精益证明智能体是形式数学推理中的核心挑战。近期工作强调围绕精益的工作流程,如智能体如何分解证明义务、使用工具和编译器反馈、诊断失败、修复证明及维护结构化证明上下文。受代码级自进化智能体启发,研究此类工作流程能否进化而非手工设计。提出一种自进化精益证明智能体,小型固定可信运行时包裹完全可变工作区。与多数针对固定外部基准优化的自进化系统不同,该系统使智能体及其基准协同进化。各代之间,得分最高的智能体通过掌握受限课程更新修改活动任务分布,并通过单锚重新校准在更新基准上重新运行冠军智能体以保持分数可比。所有进化都在基于精益的验证循环内。运行协同进化轨迹和固定基准基线15代,并在保留的miniF2F测试集上比较。最佳协同进化智能体的保留解决率达45.1%,而种子智能体为12.7%,最佳固定基准智能体为32.0%,表明基于验证器的自进化可在协同进化基准下改善精益证明工作流程。

英文摘要

Designing effective Lean proof agents is a central challenge in formal mathematical reasoning. Beyond building stronger provers, recent work emphasizes the workflow around Lean: how an agent decomposes proof obligations, uses tools and compiler feedback, diagnoses failures, repairs proofs, and maintains structured proof context. Motivated by code-level self-evolving agents, we study whether such workflows can be evolved rather than hand-designed. We present a self-evolving Lean proof agent in which a small fixed, trusted runtime wraps a fully mutable workspace: the proof workflow, prompts, and tools. Unlike most self-evolving systems, which optimize against a fixed external benchmark, our system coevolves the agent and its benchmark. Between generations, the highest-scoring agent (the champion) revises the active task distribution through a mastery-throttled curriculum update that introduces harder proof obligations only after the current level is mastered, and a single-anchor recalibration re-runs the champion on the updated benchmark to keep scores comparable as difficulty rises. All evolution stays inside a Lean-grounded verification loop: however the agent rewrites itself, a success counts only when its behavior yields Lean-verified proofs under a trusted snapshot, and each attempt must emit a machine-readable, Lean-grounded proof context whose representation may evolve but whose groundedness is enforced. We run the coevolving trajectory and a fixed-benchmark baseline for 15 active generations and compare them on a held-out miniF2F test split. The best coevolving agent reaches a 45.1% held-out solve rate, versus 12.7% for the seed and 32.0% for the best fixed-benchmark agent, showing that verifier-grounded self-evolution can improve Lean proof workflows under a coevolving benchmark.

Comments22 pages, 2 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑