arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PriorProof:形式化证明中技术新颖性的时间点度量

PriorProof: A Point-in-Time Measure of Technique Novelty for Formal Proofs

Neel Somani

arXiv 2607.16997首次发表:更新:

AI 中文总结

研究形式化数学中时间相关的证明路径非标准性,PriorProof通过提取证明项依赖足迹并评分加权惊奇度来衡量技术新颖性,无需人工标注,实验表明其与领域评分者有一定一致性,可提供可靠性指标。

AI 中文摘要

数学家能够区分解释、简化或引入非标准路径的证明,但这些判断难以实施。我们研究了形式化数学中一个更狭义的概念:时间相关的证明路径非标准性。对于一个Lean定理,PriorProof提取其详细证明项的依赖足迹,并在仅根据Mathlib的早期季度快照构建的检索条件分层平滑先验下,对该足迹的加权惊奇度进行评分。该方法无需手工构建技术本体和人工标注:语句检索从证明派生的对比对中学习,而评分对象从证明项中机械读取。在一项盲拓扑研究中,100个展示归结为76个不同的基础对:12个标准对比展示三次用于一致性筛选,64个不同的分层对。与三位保留的领域评分者中的大多数相比,PriorProof在76对中的53对(69.7%,威尔逊95%置信区间58.7 - 78.9%)上达成一致,包括12对标准对中的11对(91.7%,64.6 - 98.5%)和64对分层对中的42对(65.6%,53.4 - 76.1%)。重复归结后得分差距四分位数非单调;最小差距区间的端点为12/19(63.2%,41.0 - 80.9%),最大差距区间为16/19(84.2%,62.4 - 94.5%),支持端点校准趋势而非解决的阶梯状。最佳语言模型条件在76对中的60对(78.9%,68.5 - 86.6%)上达成一致;在配对结果上,仅PriorProof在8对中正确,仅模型在15对中正确(精确双边麦克内马尔p = 0.210),所以在这个样本量下差异未确立。因此,我们提出PriorProof不是作为专家或模型判断的替代品,而是作为一个可分解、时间锚定的信号,其得分差距提供了一个可解释的可靠性指标。

英文摘要

Mathematicians distinguish proofs that explain, simplify, or introduce a nonstandard route, but these judgments are difficult to operationalize. We study a deliberately narrower construct: time-relative proof-route nonstandardness in formal mathematics. For a Lean theorem, PriorProof extracts the dependency footprint of its elaborated proof term and scores the weighted surprisal of that footprint under a retrieval-conditioned, hierarchically smoothed prior built only from an earlier quarterly snapshot of Mathlib. The method requires no hand-built technique ontology and no human labels: statement retrieval is learned from proof-derived contrastive pairs, while the scored object is read mechanically from proof terms. In a blinded topology study, 100 presentations collapse to 76 distinct underlying pairs: 12 canonical contrasts shown three times for consistency screening and 64 distinct stratified pairs. Against the majority of three retained domain raters, PriorProof agrees on 53/76 pairs (69.7%, Wilson 95% CI 58.7-78.9%), including 11/12 canonical pairs (91.7%, 64.6-98.5%) and 42/64 stratified pairs (65.6%, 53.4-76.1%). Score-gap quartiles are nonmonotone after repeat collapse; the endpoints are 12/19 (63.2%, 41.0-80.9%) in the smallest-gap bin and 16/19 (84.2%, 62.4-94.5%) in the largest, supporting an endpoint-calibration tendency rather than a resolved staircase. The best language-model condition agrees on 60/76 pairs (78.9%, 68.5-86.6%); on paired outcomes, PriorProof alone is correct on 8 pairs and the model alone on 15 (exact two-sided McNemar p = 0.210), so the difference is not established at this sample size. We therefore present PriorProof not as a replacement for expert or model judgment, but as a decomposable, time-anchored signal whose score gap provides an interpretable reliability indicator.

Comments17 pages, 2 figures. Code: https://github.com/neelsomani/priorproof

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑