arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39228cs.AI

Fyan:一种带语义审计的人机协同框架,用于文档级形式化

Fyan: A Human--AI Harness with Semantic Auditing for Document-Level Formalization

Wei Zhao, Yangshuo Zou, Chengxiang Ding, Yifan Wu, Xuchuan Wang, Zimu Mao, Lei Zhang, Tao Luo

首次发表
浏览论文内容

中文总结 AI 辅助

FYAN是一种人机协同框架,通过证据驱动的语义审计和端到端工作流,显著提升了文档级数学形式化的准确性和一致性,并构建了大型Lean库。

中文摘要 AI 辅助

我们提出了FYAN,一种用于文档级数学形式化的人机协同框架。该框架不是孤立地处理定理,而是协调一个端到端的工作流程,涵盖规范说明、证明规划、逻辑审查、Lean证明构建、知识整理和验证,并支持独立监督和人工指导。其核心组成部分是基于证据的语义审计,用于评估形式化陈述是否忠实地保留了其非形式化规范。一个语言模型在局部对应关系、遗漏、范围和逻辑关系上构建结构化证据,而一个确定性验证器检查这些证据并产生可复现的判断。当接受一个实质性但可允许的偏差时,FYAN要求一个明确的证明转移义务,将形式化陈述与面向来源的解释联系起来。在每一阶段使用相同的模型(DeepSeek-V4.1-Flash),FYAN在严格的Lean检查下证明了143个FormalTCS定理中的86个,而通用智能体框架为69个,并将自然语言证明得分从0.501提高到0.851。在ConsistencyCheck上,其语义审计比直接LLM评判者捕捉到更多不一致的陈述,无论是在针对来源验证的标签上(召回率0.777对0.636)还是在原始标签上(0.873对0.820),并将其报告的每个不匹配定位到特定的假设、结论或范围。FYAN还构建了ODENumLib,一个用于常微分方程数值分析的9,355行Lean库。

英文摘要

We present FYAN, a human--AI harness for document-level mathematical formalization. Rather than treating theorems in isolation, FYAN coordinates an end-to-end workflow spanning specification, proof planning, logical review, Lean proof construction, knowledge curation, and validation, with support for independent supervision and human guidance. A central component is evidence-grounded semantic auditing, which assesses whether formal statements faithfully preserve their informal specifications. A language model constructs structured evidence over local correspondences, omissions, scope, and logical relations, while a deterministic validator checks this evidence and produces reproducible judgments. When a substantive but admissible deviation is accepted, FYAN requires an explicit proof-transfer obligation connecting the formal statement back to a source-facing interpretation. With the same model (DeepSeek-V4.1-Flash) in every stage, FYAN proves 86 of 143 FormalTCS theorems under a strict Lean check, against 69 for a general agent harness, and raises the natural-language proof score from 0.501 to 0.851. On ConsistencyCheck, its semantic audit catches more inconsistent statements than a direct LLM judge, both on labels verified against the source (recall 0.777 vs. 0.636) and on the original labels (0.873 vs. 0.820), and localizes each mismatch it reports to a specific hypothesis, conclusion, or scope. FYAN also built ODENumLib, a 9,355-line Lean library for the numerical analysis of ordinary differential equation.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • University of California, Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

↑