发表机构
University of Maryland; Georgia Institute of Technology; Cornell University; All Purpose AI(马里兰大学; 佐治亚理工学院; 康奈尔大学; All Purpose AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出涵盖13个数学领域的诊断基准MathAdv,结合Lean 4评估当代定理证明器,发现形式化是瓶颈、性能跨领域差异大等问题,为分组件评估模型能力提供依据。
AI 中文摘要
形式化定理证明可实现数学推理的机器可验证评估,但现有基准常侧重整体证明准确率,聚焦于狭窄的数学领域,且仅提供有限的等价重述鲁棒性证据。我们推出MathAdv,这一诊断基准涵盖本科与研究生级数学的13个领域。结合Lean 4定理证明器,MathAdv提供最多三项辅助任务:探查数学知识的选择题、分离非形式化推理的填空题,以及测试问题呈现鲁棒性的专家精心设计的变换题。我们对当代定理证明器的评估得出四项发现:形式化仍是主要瓶颈;不同数学领域的性能差异显著;自然语言指导对通用大语言模型(LLMs)有帮助,但可能阻碍证明专用模型;数学上等价的重述暴露出严重的鲁棒性局限。这些结果共同表明,分组件评估如何能揭示整体定理证明准确率所掩盖的模型能力与失败模式。该数据集与评估脚本可在指定URL获取。
英文摘要
Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of robustness to equivalent reformulations. We introduce MathAdv, a diagnostic benchmark spanning 13 domains across undergraduate- and graduate-level mathematics. Alongside Lean 4 theorem proving, MathAdv provides up to three auxiliary tasks: multiple-choice questions that probe mathematical knowledge, fill-in-the-blank problems that isolate informal reasoning, and expert-crafted transformations that test robustness to problem presentation. Our evaluation of contemporary theorem provers yields four findings: formalization remains a major bottleneck; performance varies substantially across mathematical domains; natural-language guidance helps general-purpose LLMs but can hinder proof-specialized models; and mathematically equivalent reformulations expose substantial robustness limitations. Together, these results show how component-wise evaluation can reveal model capabilities and failure modes that aggregate theorem-proving accuracy obscures. The dataset and evaluation scripts are available at https://github.com/margotyjx/MathAdv.git.