arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Navier-Stokes 在翻译中迷失:为何对 AI 自动形式化的 Lean 验证不能保证自然语言证明的正确性

Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs

Alexander Bastounis, Fabian Circelli, Anders C. Hansen

arXiv 2610.08144首次发表:更新:

发表机构

University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文指出 AI 自动形式化中 Lean 验证无法保证自然语言证明的正确性,因语义忠实翻译的难度在 SCI 层级中为无穷,高于停机问题,并以 OpenAI 的 Navier-Stokes 证明为例说明不匹配。

AI 中文摘要

自动形式化(autoformalisation)越来越多地被用于验证数学文本,包括由 AI 生成的文本,例如 OpenAI 宣布的关于 Navier-Stokes 方程解爆破的证明。在此过程中,AI 系统将文本从自然语言(NL)翻译成形式语言(如 Lean)。一旦翻译完成,形式语言表达的论证可以轻松地被机械验证。本文旨在论证为何此过程可能无法为原始 NL 论证提供任何信心,原因在于进行语义忠实的翻译存在各种困难。特别地,我们强调,解决数学 NL 文本中的歧义问题(这是提供语义忠实翻译所必需的)在可解性复杂度指数(SCI)层级/算术层级中任意高(SCI = ∞)。因此,非正式地说,提供语义忠实的 AI 自动形式化比任何计算问题(包括停机问题,其 SCI = 1)都更难。为了展示此结果的影响,我们提供了几个 AI 将 NL 陈述和证明误译为 Lean 的实际例子,导致 NL 证明与其 Lean“验证”之间出现不匹配。这些例子包括 OpenAI 宣布的 Navier-Stokes 证明。特别地,我们表明形式化的 Lean 证明并不对应于 NL 中关于 Navier-Stokes 方程解爆破的证明。

英文摘要

Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI's announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully. In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation, is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy (the SCI $= \infty$). Hence, informally, providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI $= 1$). To demonstrate the effect of this result we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean `verifications'. These include OpenAI's announced Navier-Stokes proof. In particular, we show that the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations.

Comments25 pages, 4 Figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑