Magenta:弥合数学推理与Lean验证之间的循环
Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification
浏览论文内容
中文总结 AI 辅助
Magenta是一个无需训练的智能体流水线,将Lean验证信号融入非正式推理,在多个奥林匹克数学基准上达到100%准确率,并解决了全部IMO 2026问题。
中文摘要 AI 辅助
大多数数学知识是通过所谓的非正式数学使用和自然语言来交流的。由于大型语言模型(LLMs)高度擅长使用自然语言,它们在非正式数学推理中表现出强大的性能,但并非完美。将LLMs限制在非正式推理中,会错失利用机器通过机器可检查证明所提供的离散验证能力的机会。在本文中,我们通过将Lean信号整合到非正式推理过程中,弥合了非正式与正式推理之间的差距。我们引入了Magenta,一个无需训练的主体化流水线,仅给定一个自然语言问题,即可生成答案,将其表达为Lean 4语句,并构造机器检查的证明。一个语句评判器验证形式化是否保留了原始问题,而一个错误归因评判器将失败的尝试路由到数学重新推导或局部Lean修复。Magenta在所有评估的奥林匹克基准测试中达到了100%的准确率,包括AIME 2025、AIME 2026和HMMT 2026年2月。当与开放权重的K2-Horizon-7B推理器配对时,它解决了所有六个IMO 2026问题。我们的分析表明,语句裁决对于防止虚假证书至关重要,并且反馈引导的纠正优于在困难问题上的独立重采样。
英文摘要
Most of mathematical knowledge has been communicated through so-called informal use of mathematics and natural language. With large language models (LLMs) being highly adept in using natural language, they achieve strong performance, yet not perfect, in informal mathematical reasoning. Restraining LLMs to informal reasoning misses out on the opportunity to use the discrete verification abilities that machines offer through machine-checkable proofs. In this paper, we bridge the gap between informal and formal reasoning by integrating Lean signals into the informal reasoning process. We introduce Magenta, a training-free agentic pipeline that, given only a natural-language problem, produces an answer, expresses it as a Lean 4 statement, and constructs a machine-checked proof. A statement judge verifies whether the formalisation preserves the original problem, while an error-attribution judge routes failed attempts either to mathematical re-derivation or local Lean repair. Magenta achieves 100% accuracy across all evaluated olympiad benchmarks, including AIME 2025, AIME 2026, and HMMT February 2026. When paired with the open-weight K2-Horizon-7B reasoner, it solves all six IMO 2026 problems. Our analysis shows that statement adjudication is essential for preventing false certificates and that feedback-guided correction outperforms independent resampling on difficult problems.
发表机构
- Institute of Foundation Models(基础模型研究所)
- Imperial College London(伦敦帝国理工学院)
- University of Edinburgh(爱丁堡大学)
- University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。