发表机构
Brown University; California Institute of Technology(布朗大学; 加州理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究形式定理证明生态系统分散问题,提出ITPEval基准测试,涵盖四个主要交互式定理证明器及两种逻辑基础,评估语句和证明翻译,揭示库不匹配是瓶颈,发布相关基础设施和管道,展示多方面评估结果。
AI 中文摘要
形式定理证明已成为机器学习的前沿挑战,但生态系统分散,证明在不兼容系统中孤立,限制了基于学习的证明器的训练数据和验证结果的可移植性。我们提出了ITPEval,这是首个用于评估四个主要交互式定理证明器(Lean 4、Rocq、Isabelle和HOL Light)之间自动形式证明翻译的基准测试,涵盖两种不同逻辑基础。我们的基准包括1560个源文件和6848个定理,分为隔离基础翻译难度的公理化文件控制层和暴露API及证明风格不匹配的真实库生态系统层。我们发布了itpeval,这是一个统一的多交互式定理证明器验证基础设施,具有状态隔离的热后端,保留每个工件的原生检查语义。我们在12个定向翻译对上评估了五个前沿和开放权重的语言模型的语句和证明翻译:语句翻译的通过率@1峰值为29.1%,证明翻译为10.5%;控制定理的证明通过率@1达到29.7%,而生态系统级翻译为%,证实库不匹配是主要瓶颈。除了通过率@k评估外,确定性的Lean 4 BEq检查为54.0%的经过验证的源到Lean 4 miniF2F语句翻译建立了等价性,表明仅原生类型检查可能会高估语义保真度;在自动形式化/自动非形式化往返研究中,Rocq和HOL Light比Lean 4和Isabelle更容易形式化,而多交互式定理证明器上下文将Lean 4的合并成功率从4.8%提高到10.6%。我们的基准、验证基础设施和评估管道已公开发布。
英文摘要
Formal theorem proving has emerged as a frontier challenge for machine learning, yet the ecosystem is fragmented: proofs remain siloed across incompatible systems, limiting both training data for learning-based provers and the portability of verified results. We present ITPEval, the first benchmark for evaluating automated formal proof translation across four major ITPs (Lean 4, Rocq, Isabelle, and HOL Light), spanning two distinct logical foundations. Our benchmark comprises 1,560 source files and 6,848 theorems organized into a controlled tier of axiomatized files that isolates foundational translation difficulty, and an ecosystem tier drawn from real libraries that exposes API and proof-style mismatches. We release itpeval, a unified multi-ITP verification infrastructure with state-isolated warm backends that preserve per-artifact native checking semantics. We evaluate both statement and proof translation across five frontier and open-weight LLMs on 12 directed translation pairs: statement translation peaks at 29.1% pass@1 and proof translation at 10.5%; controlled theorems reach 29.7% proof pass@1 versus 5.2% for ecosystem-level translations, confirming that library mismatch is the dominant bottleneck. In addition to pass@k evaluation, a deterministic Lean 4 BEq check establishes equivalence for 54.0% of verified source-to-Lean 4 miniF2F statement translations, showing that native type-checking alone can substantially overestimate semantic fidelity; in an autoformalization/auto-informalization round-trip study, Rocq and HOL Light are easier formalization targets than Lean 4 and Isabelle, while multi-ITP context improves pooled Lean 4 success from 4.8% to 10.6%. Our benchmark, verification infrastructure, and evaluation pipelines are publicly released.
Comments23 pages