arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SCOPE:以语言模型作为策略规划器的认证定理证明

SCOPE: Certified Theorem Proving with a Language Model as the Policy Planner

Hanchao Zhou, Jialei Li

arXiv 2610.08319首次发表:更新:

AI 中文总结

SCOPE通过让语言模型在算子词汇上规划、符号引擎执行数值、编译器生成证明,在218题测试集上以135M模型认证191题(87.6%),远超更大模型,验证了分工式证明生成的有效性。

AI 中文摘要

在Lean等证明助手中,生成的证明必须通过机器编译检查,因此评估无需人工评分。直接生成在多步数值命题上会失败:证明仅在每个内容整数都正确时才有效,因此通过率受限于每个整数准确率的k次幂。对2,617个参考证明进行的受控破坏实验证实了这一幂律。SCOPE(状态条件算子规划与执行)强制实施自然的劳动分工:模型在算子词汇表上进行规划,符号引擎执行数值计算,编译器生成证明。在218个问题的测试集上,它使用135M骨干模型认证了191/218(87.6%);7B规模的DeepSeek-Prover-V1.5-RL在27.5倍的token数和37.5倍的墙钟时间下仅认证了18/218,而DeepSeek-Prover-V2-7B在双向双测试集上认证数为零。多步思考每个问题花费6.12个离散决策动作,且不产生自然语言思考文本。将决策框架中的滞后引擎状态替换为当前状态,使通过率从117/218提升至191/218,而增加链末端损失的权重则有害。在公开的Lean-Workbook库中,3,536个可评分可接受问题中有2,132个获得认证(60.29%),且主测试集上零回归。所有读数均来自版本冻结的评审,并进行了独立复核和反向验证。限制自由生成并保持决策时信息可见,是比扩大模型规模更直接的途径。

英文摘要

In proof assistants such as Lean, a generated proof must pass machine compilation checks, so evaluation needs no human scoring. Direct generation fails on multi-step numeric propositions: a proof is valid only if every content integer is correct, so the pass rate is bounded by the k-th power of the per-integer accuracy. Controlled corruption across 2,617 reference proofs confirms this power law. SCOPE (State-Conditioned Operator Planning and Execution) enforces the natural division of labor: the model plans over an operator vocabulary, a symbolic engine executes the numerics, and a compiler renders the proof. On a 218-problem suite it certifies 191/218 (87.6%) with a 135M backbone; the 7B DeepSeek-Prover-V1.5-RL certifies 18/218 at 27.5 times the tokens and 37.5 times the wall-clock, and DeepSeek-Prover-V2-7B certifies zero on a bidirectional dual suite. Multi-step thinking costs 6.12 discrete decision actions per problem and produces no natural-language thinking text. Replacing the lagged engine state in the decision frame with the current one lifts the pass rate from 117/218 to 191/218, while up-weighting the chain-end loss hurts. On the public Lean-Workbook library, 2,132 of 3,536 gradeable admissible problems certify (60.29%) with zero regression on the main suite. All readings come from a version-frozen review with independent rechecks and reverse verification. Restricting free generation and keeping decision-time information visible is a more direct route than enlarging the model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑