发表机构
University of Virginia; Meta AI(弗吉尼亚大学; Meta AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ProofEvolve是一种神经符号框架,结合神经模型与Lean内核进化形式化验证的证明结构,在三个竞赛级Lean基准中实现最高平均定理解决率,解决现有神经证明器的递归结构缺失问题。
AI 中文摘要
自动定理证明为科学发现中的递归自我改进提供了自然基础。然而,现有的神经证明器未能充分保留这种递归结构,而学习过程本应随时间实现自我改进。现有方法要么通过代价高昂的权重更新将证明经验嵌入模型参数,要么仅在当前问题中保留已验证的中间演绎。此外,这些方法还严重依赖稀疏的完整证明反馈,即便未成功的部分尝试也可能包含有用发现。为缩小这一差距,我们提出ProofEvolve,这是一种神经符号框架,它结合神经模型进化显式、经形式化验证的符号证明结构,以果断扩展知识边界。在该框架中,神经模型提出变异算子,包括分解、修复和模式重组;符号Lean内核验证每一次证明转换。在进化循环中,ProofEvolve对生成的证明有向无环图(DAG)计算已验证闭包。在每个问题内,ProofEvolve在行为索引档案中进化部分与或证明DAG;跨问题时,经内核检查的模式提取将新证明的子DAG添加到持久模式库。证明DAG通过类型化模式重组继承已解决结果,每个剩余前提作为新子目标暴露。这种进化过程保留不完整尝试的已验证结果,使其可用于后续证明,同时不削弱形式化可靠性。在三个竞赛级Lean基准测试中,ProofEvolve在评估的证明系统中实现最高平均解决率。
英文摘要
Automated theorem proving offers a natural foundation for recursive self-improvement in scientific discovery. However, existing neural provers do not fully preserve this recursive structure, where the learning process should be self-improving over time. Existing methods either embed proof experience into model parameters through expensive weight updates, or keep verified intermediate deductions only within the current problem. In addition, these methods also heavily rely on sparse whole-proof feedback, even when unsuccessful partial attempts contain useful discoveries. To close the gap, we propose ProofEvolve, a neuro-symbolic framework that evolves explicit, formally verified symbolic proof structures with neural models to decisively expand the knowledge boundary. In this framework, the neural model proposes variation operators, including decompositions, repairs, and schema recombinations. The symbolic Lean kernel verifies every proof transition. Over the evolution loops, ProofEvolve computes verified closure over the resulting proof directed acyclic graphs (DAGs). Within each problem, ProofEvolve evolves partial AND-OR proof DAGs in a behaviorally indexed archive. Across problems, kernel-checked schema extraction adds newly proved sub-DAGs to a persistent schema library. Proof DAGs inherit the solved results through typed schema recombination, with every residual premise exposed as a new subgoal. This evolutionary process preserves verified results from incomplete attempts and makes them available for later proofs without weakening formal soundness. Across three competition-level Lean benchmarks, ProofEvolve achieves the highest average solve rate among the evaluated proof systems.