arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20432cs.LOcs.AIcs.CL

ProofJudge:基于工具的大语言模型(LLM)对Mathlib中形式化证明质量的评估工具

ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib

  • Dreadnode(德雷德节点)

机构由 AI 辅助整理,请以论文原文为准。

Shane Caldwell

AI总结:

本研究提出ProofJudge系统,通过智能体式LLM从五个维度评估Mathlib形式化证明质量,在218个声明数据集上验证其与人类偏好对齐度,发布开源制品支持相关研究。

AI中文摘要:

通过Lean 4内核类型检查器的形式化证明,其质量仍可能存在巨大差异。我们提出ProofJudge,这是一种智能体式的大语言模型(LLM)作为评判者的系统,它在正确性之外,从五个维度对形式化证明质量进行评分:库利用度、适配自动化工具的程度、结构清晰度、命题陈述质量以及Mathlib规范。我们从不同的Mathlib拉取请求(PR)中选取218个声明,构建了一个新的数据集,并在该数据集上对ProofJudge进行评估。该评判智能体通过工具访问PR所应用的提交内容,从而在评分时能够查询库的状态。当评判智能体将Mathlib所接受的PR版本评为高于被退回修改的初始版本时,即认为其与人类偏好对齐。所有6个被评估的评判模型恢复评审人员偏好的准确率均远高于随机水平,从80.8%到63.5%不等,且两个开放权重的评判模型的准确率约为70%,而成本仅为最佳评判模型的十分之一。我们将评判框架、评估数据集和评估轨迹作为开源制品发布,以支持进一步的研究。

英文摘要:

Formal proofs in Lean 4 that pass the kernel's type checker can nonetheless vary widely in quality. We introduce ProofJudge, an agentic LLM-as-judge system that scores formal proof quality along five dimensions beyond correctness: library leverage, automation fit, structural clarity, statement quality, and Mathlib conventions. We evaluate ProofJudge on a novel dataset of 218 declarations drawn from distinct Mathlib PRs. The judge agent is grounded by tool access to the commit the PR is applied to, enabling it to query the library state when scoring. A judge is considered aligned with human preferences when it rates the version of the PR Mathlib accepted above the initial version that was sent back for revision. All six judge models evaluated recover the reviewers' preference well above chance, from 80.8% to 63.5%, and two open-weight judges reach roughly 70% at a tenth of the best judge's cost. We release the judge harness, evaluation dataset, and evaluation traces as open-source artifacts to support further research.

补充信息

↑