发表机构
Universidad Nacional del Centro de la Provincia de Buenos Aires (UNICEN)(布宜诺斯艾利斯省国立中部大学(UNICEN))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 AViD Journal 流程,基于 Lean 4 从三维度验证数学定理新颖性,评估时发现编译不保证语义保真度等三个通用障碍。
AI 中文摘要
应用于数学的人工智能系统仅能验证正确性,无法验证新颖性:自动生成的定理可在 Lean 4 中编译无错,但仍可能是已有的已知结果。本文提出 AViD Journal,这是一个接收 LaTeX 文章、在 Lean 4 中形式化其陈述、并通过三个维度的决策树给出新颖性判定的流程:形式化语料库(Mathlib)和非形式化语料库(TheoremSearch 与 Matlas,带时间过滤器和大语言模型判定)中的先验存在性、通过自动战术实现的非平凡性、以前提集的雅卡尔距离衡量的证明结构距离。对因声明重复而从 arXiv 撤回的论文的评估得出了比任何性能指标更具信息量的结果:确定了三个无论该实现如何都会限制方法的障碍。第一,Lean 文件的成功编译不保证语义保真度。第二,召回上限由定理索引的覆盖范围决定,而非相似度度量。第三,arXiv 在论文撤回时会删除其源代码,损害基于这些论文构建的任何基准的可复现性。
英文摘要
Artificial intelligence systems applied to mathematics verify correctness but not novelty: an automatically generated theorem can compile in Lean without errors and yet be an already known result. This article presents AViD Journal, a pipeline that receives a LaTeX article, formalizes its statements in Lean 4, and issues a novelty verdict through a decision tree over three dimensions: prior existence in a formal corpus (Mathlib) and an informal one (TheoremSearch and Matlas, with temporal filter and LLM judge), non-triviality via automatic tactics, and structural distance between proofs measured as Jaccard distance over premise sets. Evaluation on papers withdrawn from arXiv due to declared duplication produced a result more informative than any performance measure: the identification of three obstacles that limit the approach regardless of this implementation. First, successful compilation of a Lean file does not guarantee semantic fidelity. Second, the recall ceiling is imposed by the coverage of theorem indices, not by the similarity metric. Third, arXiv removes the source code of articles upon withdrawal, compromising the reproducibility of any benchmark built upon them.
Comments20 pages. Preliminary version; a large-scale quantitative evaluation (N theorems x M models) is left to future work. Comments welcome. Code: https://github.com/ayrtonporto/avid-journal