arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MathForm:结合知识检索与验证引导的迭代优化实现数学自动形式化的规模化

MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement

Lushi Pu, Weiming Zhang, Xinheng Xie, Zixuan Fu, Bingxiang He, Hengyu Zhao, Hongya Lyu, Xin Li, Jie Zhou, Yudong Wang

arXiv 2608.14221首次发表:更新:

发表机构

ModelBest Inc.; Tsinghua University(ModelBest公司; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出MathForm框架,结合Mathlib知识检索与验证引导迭代优化构建Lean 4数据集,训练的MathForm-8B在多项基准测试中超越同类模型,实现了数学自动形式化的性能提升。

AI 中文摘要

自动形式化通常被定义为将自然语言数学语句翻译为Lean 4等机器可验证的形式语言,但忠实的形式化不仅需要翻译,模型还需将数学概念映射到Mathlib等形式库的复杂类型与定义层级,同时确保生成语句保留源命题的含义。现有方法存在不足,因为它们严重依赖模型的参数记忆来存储库特定知识,而常见的数据构建流程往往仅过滤单次生成结果,缺乏反馈驱动的修订机制。为应对这些挑战,我们提出MathForm,这是一个通过Mathlib知识检索和验证引导的迭代优化构建已验证训练数据的自动形式化框架。生成前,检索规划器从Mathlib收集相关定义与现有形式化结果,以引导形式化生成器;生成的语句随后通过编译器诊断和语义一致性反馈进行修订。利用该框架,我们构建了FormalVerse,这是一个包含约36.7万个跨不同数学领域与来源的已验证示例的Lean 4数据集。接着,我们通过监督微调后再进行强化学习训练MathForm-8B。在六个基准测试中,MathForm-8B在语法检查(SC)下达到平均Pass@8率88.06%,在一致性检查(CC)下达到72.37%,表现优于多个专用的32B规模自动形式化模型;在具有挑战性的FATE-H和FATE-X子集上,其CC通过率分别达到63%和37%,在这两个子集上均超过了最强的专用基线模型。

英文摘要

Autoformalization is commonly framed as translating natural-language mathematical statements into machine-verifiable formal languages such as Lean 4. However, faithful formalization requires more than translation. Models must map mathematical concepts to the complex hierarchy of types and definitions in formal libraries such as Mathlib, while ensuring that generated statements preserve the meaning of the source propositions. Existing approaches struggle because they rely heavily on the model's parametric memory for library-specific knowledge, while common data construction pipelines often resort to filtering single-pass outputs and lack mechanisms for feedback-driven revision. To address these challenges, we introduce MathForm, an autoformalization framework for constructing verified training data through Mathlib knowledge retrieval and verification-guided iterative refinement. Before generation, a retrieval planner gathers relevant definitions and existing formalizations from Mathlib to guide the formalization generator. Generated statements are then revised using compiler diagnostics and semantic-consistency feedback. Using this framework, we construct FormalVerse, a Lean 4 dataset containing approximately 367K verified examples across diverse mathematical domains and sources. We then train MathForm-8B through supervised fine-tuning followed by reinforcement learning. Across six benchmarks, MathForm-8B achieves average Pass@8 rates of 88.06% under Syntax Check (SC) and 72.37% under Consistency Check (CC), outperforming multiple specialized 32B autoformalizers. On the challenging FATE-H and FATE-X subsets, it attains CC pass rates of 63% and 37%, exceeding the strongest specialized baselines in both cases.

Comments25 pages, 6 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑