arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AIProver:通过证书驱动的演化工具实现数学研究的智能体自动形式化

AIProver: Agentic Auto-Formalization of Mathematical Research via Certificate-Driven Evolving Harness

Prithwish Jana, Viet Bach Hoang, Logan Luna, Viresh Pati, Akash Singirikonda, Cy Xie, Lisa Carbone, Wuyang Chen, Walter Moreira, Joe Stubbs, Sriram Vishwanath, Vijay Ganesh

arXiv 2610.05367首次发表:更新:

发表机构

Georgia Institute of Technology; University of Pennsylvania; Foothill College; Rutgers University; Simon Fraser University; University of Texas at Austin(佐治亚理工学院; 宾夕法尼亚大学; 山麓学院; 罗格斯大学; 西蒙弗雷泽大学; 德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AIProver 通过证书驱动的 HarnessEvolve 演化工具,联合后训练 119B 模型,在 LoCoBench 上实现研究级证明自动形式化,显著提升语义正确性并降低成本。

AI 中文摘要

证明自动形式化将自然语言(NL)定理和证明转换为形式语言(FL),例如 Lean,从而实现机械验证。尽管进展迅速,但研究级别的证明通常依赖于领先的证明助手库(例如 Lean 的 Mathlib)中缺失的概念,并且成功编译并不能保证翻译保留定理的含义或证明的推理过程。此外,对齐的 NL-FL 训练数据稀缺,领先的智能体通常依赖昂贵的前沿模型和手动设计的工具。为了解决上述问题,我们提出了 AIProver,一个用于自主证明自动形式化和证明合成(AFPS)的智能体框架,该框架联合后训练了一个 119B 开放权重语言模型,并通过 HarnessEvolve 演化其智能体的、工具调用的工具。验证器评估类型正确性、证明完整性和语义正确性,返回奖励和诊断证书,这些证书驱动模型微调和通过符号反馈进行的交替强化学习,以及 HarnessEvolve,一种证书驱动的对整个工具控制流的演化搜索,该搜索根据更新后的模型重新定制工具。为了进行研究级别的训练和评估,我们引入了 LoCoBench,包含来自 Mathlib、CSLib、Mizar 数学库和一本有界算术教科书的 58.9k 个实例,并有一个 771 个实例的验证集,其定理-证明对没有公开的 Lean 形式化。与涵盖 AFPS 智能体、前沿 LLM 和编码智能体的 39 个框架相比,AIProver 将其 pass@4 语义正确性从其 Leanstral-1.5 基础从 15.7% 提升到 36.7%,并优于所有其他开放权重系统和 Aristotle。作为 Claude Code 和 Codex 技能,它将其语义正确性分别从 41.9% 和 34.1% 提升到 79.8% 和 62.4%。此外,它比 Numina-Lean-Agent 便宜 24%,推动了研究级 AFPS 的精度-成本前沿。

英文摘要

Proof auto-formalization translates natural-language (NL) theorems and proofs into a formal language (FL) such as Lean, enabling mechanical verification. Despite rapid progress, research-level proofs often depend on concepts missing from leading proof assistant libraries (e.g., Lean's Mathlib), and successful compilation does not guarantee that a translation preserves the theorem's meaning or the proof's reasoning. Furthermore, aligned NL-FL training data are scarce, and leading agents often rely on costly frontier models and manually engineered harnesses. To address the above issues, we present AIProver, an agentic framework for autonomous proof auto-formalization and proof synthesis (AFPS) that jointly post-trains a 119B open-weight language model and evolves its agentic, tool-calling harness with HarnessEvolve. Verifiers assess type correctness, proof completeness, and semantic correctness, returning rewards and diagnostic certificates that drive model fine-tuning and alternating reinforcement learning via symbolic feedback and HarnessEvolve, a certificate-driven evolutionary search over the whole harness control flow that re-tailors the harness to the updated model. For research-level training and evaluation, we introduce LoCoBench, 58.9k instances from Mathlib, CSLib, Mizar Math Library, and a bounded-arithmetic textbook, with a 771-instance validation split whose theorem-proof pairs have no public Lean formalization. Against 39 frameworks spanning AFPS agents, frontier LLMs, and coding agents, AIProver lifts pass@4 semantic correctness over its Leanstral-1.5 base from 15.7% to 36.7% and outperforms every other open-weight system and Aristotle. As a Claude Code and Codex skill, it lifts their semantic correctness from 41.9% and 34.1% to 79.8% and 62.4%, respectively. Further, it is also 24% cheaper than Numina-Lean-Agent, pushing the accuracy-cost frontier of research-level AFPS.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑