发表机构
Google DeepMind; ByteDance Seed Team(谷歌DeepMind; 字节跳动种子团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对几何问题形式化的碎片化现状,提出Euclean框架,通过四个阶段在Mathlib中自动形式化几何,构建了大型数据集,经评估有一定准确率,能提升Goedel v2证明成功率,验证了数据集质量。
AI 中文摘要
近期形式推理系统已达国际数学奥林匹克竞赛水平,但存在碎片化问题:代数和数论在Lean中处理,几何仍依赖形式保证有限的特定领域语言。这增加了可信计算基并阻碍统一模型开发。现有Lean中的几何研究引入与标准Mathlib不兼容的自定义公理系统且规模小。本文提出Euclean框架,构建了大型几何形式化数据集,经人类评估有一定准确率,验证了数据集对统一神经定理证明的质量。
英文摘要
Recent formal reasoning systems have reached IMO-level performance, yet they leave a fragmented landscape: algebra and number theory are handled in Lean, while geometry still relies on domain-specific languages with limited formal guarantees. This split increases the trusted computing base and hinders unified model development. Existing geometry-in-Lean efforts (LeanEuclid, LeanGeo) introduce custom axiom systems incompatible with standard Mathlib, and their small scale ($<$ 1,100 problems) limits large-scale training. Native Mathlib autoformalization of geometry, however, poses distinct challenges: implicit diagrammatic assumptions (e.g., topological configuration and non-degeneracy) must be made explicit rather than deferred to external solvers, and models must adapt to Mathlib's small, rapidly evolving geometry infrastructure. We present Euclean, a four-stage framework - constraint explication, configuration anchoring, formalization mapping, and iterative repair - for automatically formalizing geometry in native Mathlib. We construct OMNI-Geometry (768 competition problems) and Numina-Geometry (177,597 problems), the largest geometry formalization dataset in Lean. Human evaluation shows 48.89% TOP1 and 73.33% TOP5 accuracy. Training Goedel v2 on our formalizations improves proof success from 13.6% to 15.1%, validating dataset quality for unified neural theorem proving. Code and datasets: https://github.com/tlb-22/Euclean.
CommentsICML 2026