发表机构
Moonshot AI; epfl-lara(月球计划人工智能公司; 洛桑联邦理工学院拉腊实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究通过对两篇数学论文案例研究,探讨文档到项目形式化中运行时机制对相关指标的影响,使用Kimi2.6和GPT5.5进行消融实验,报告多项数据,展示LeanFlow在特定任务上的表现及校准成果。
AI 中文摘要
我们展示并评估了LeanFlow,这是一个专门用于将数学论文翻译成可构建的Lean项目的大语言模型代理系统。最近的循环验证系统表明可以生成大型形式工件,但在文档到项目的形式化中,尚不清楚哪些运行时机制会影响完成度、可审计性或效率。我们通过对两篇数论和测度论中先前未形式化的数学论文进行案例研究来探讨这个问题,使用Kimi2.6和GPT5.5进行模型、证明工作流和工具集消融;我们报告任务结果、API调用、输入令牌和输出令牌。使用Kimi2.6时,完整工作流在2000次调用预算内完成了两个文档级项目,而无队列变体达到了预算限制;使用GPT5.5时,所有文档级变体都完成了,并且完整工作流在两个来源上的输入令牌成本最低或并列最低。作为补充校准,LeanFlow在RLM25的PFR切片上达到了75.7%的BEq+,并在我们的GPT5.5运行中解决了所有五个ICML 2026数学人工智能TCS挑战项目。
英文摘要
We present and evaluate LeanFlow, an LLM agent system specialized for translating mathematical papers into buildable Lean projects. Recent verifier-in-the-loop systems show that large formal artifacts can be produced, but it remains unclear which runtime mechanisms affect completion, auditability, or efficiency in document-to-project formalization. We study this question through case studies on two previously unformalized mathematical papers in number theory and measure theory, using model, proof-workflow, and toolset ablations with Kimi2.6 and GPT5.5; we report task outcome, API calls, input tokens, and output tokens. With Kimi2.6, the full workflow completes both document-level projects within the 2000-call budget, while no-queue variants reach the budget limit; with GPT5.5, all document-level variants complete, and the full workflow has the lowest or tied-lowest input-token cost on both sources. As complementary calibration, LeanFlow reaches 75.7% BEq+ on the PFR slice of RLM25 and solves all five ICML 2026 AI for Math TCS challenge projects in our GPT5.5 runs.
Comments14 pages, 3 figures, ICML 2026: AI for Math Workshop