arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35790cs.LGcs.AIcs.CLcs.LO

Sage:带语义校正的形式化

Sage: Formalization with Semantic Correction

  • Huawei Lagrange Mathematics Computing Research Center(华为拉格朗日数学计算研究中心)

机构由 AI 辅助整理,请以论文原文为准。

Thomas Hirtz, Farzad Jafarrahmani, Abdelmouksit Sagueni, Xiang Zhou, Wenping Deng, Liang Zhang

AI总结:

Sage 是一个智能体框架,通过四阶段分解生成与双信号语义校正循环,解决自然语言到 Lean 4 形式化的“严谨性幻觉”问题,显著降低答案泄漏并提升编译与语义保真度。

AI中文摘要:

尽管神经定理证明器在形式数学领域取得了令人瞩目的里程碑式进展,但它们大多基于一个假设:即已经提供了忠实的 Lean 4 形式化陈述。将非正式的自然语言翻译成形式语言是一个关键的数据瓶颈,且受到“严谨性幻觉”的困扰:标准类型检查器会接受那些能够编译但会丢失假设、引入空洞真值或微妙地改变数学界限的陈述。为解决这一问题,我们引入了 Sage(语义智能体引导形式化引擎),这是一个智能体框架,它将整体式翻译替换为四阶段分解生成流水线,并配备双信号语义校正循环。通过将 Lean 4 编译器诊断与多维语义反馈相结合,我们的校正循环在确保句法有效性的同时强化了数学保真度。通过显式考虑开放式查询与声明式形式目标之间的差距,我们的流水线防止了模型通过猜测未经验证的答案来获得高形式化率(表现出 70.9% 的答案泄漏率)。因此,Sage 将泄漏率抑制到 2.7%,同时在 Omni-MATH 无证明数据集上实现了 73.3% 的 pass@4 联合编译与语义保真度(相比之下,微调的 Goedel-Formalizer-V2 基线为 42.0%)。最后,在 IMO-Unformalized(一个包含 175 个未形式化的国际数学奥林匹克问题的新前沿基准)上,Sage 展示了有效的零样本泛化能力,其 pass@4 验证保真度达到 87.4%,而基线仅为 19.4%,并在超过 79% 的盲法成对评估中获胜。

英文摘要:

While neural theorem provers have achieved impressive milestones in formal mathematics, they largely operate on the assumption that faithful Lean 4 formal statements are already provided. Translating informal natural language into a formal language is a critical data bottleneck plagued by an "illusion of rigor": standard type-checkers accept statements that compile but drop hypotheses, introduce vacuous truths, or subtly alter mathematical bounds. To resolve this, we introduce Sage (Semantic Agent-Guided Formalization Engine), an agentic framework that replaces monolithic translation with a four-stage decomposed generation pipeline coupled with a dual-signal semantic correction loop. By pairing Lean 4 compiler diagnostics with multi-dimensional semantic feedback, our correction loop enforces mathematical fidelity alongside syntactic validity. By explicitly accounting for the gap between open-ended queries and declarative formal targets, our pipeline prevents models from achieving high formalization rates by guessing unverified answers (exhibiting a 70.9% answer leakage rate in monolithic baselines). Consequently, Sage suppresses leakage to 2.7% while achieving 73.3% pass@4 joint compilation and semantic fidelity on the Omni-MATH without proofs (compared to 42.0% for a fine-tuned Goedel-Formalizer-V2 baseline). Finally, on IMO-Unformalized, a novel frontier of 175 unformalized International Mathematical Olympiad problems, Sage demonstrates effective zero-shot generalization with 87.4% pass@4 verified fidelity compared to just 19.4% for the baseline, winning over 79% of blind pairwise evaluations.

补充信息

↑