Fyan:一种带语义审计的人机协同框架,用于文档级形式化
Fyan: A Human--AI Harness with Semantic Auditing for Document-Level Formalization
浏览论文内容
中文总结 AI 辅助
FYAN是一种人机协同框架,通过证据驱动的语义审计和端到端工作流,显著提升了文档级数学形式化的准确性和一致性,并构建了大型Lean库。
中文摘要 AI 辅助
我们提出了FYAN,一种用于文档级数学形式化的人机协同框架。该框架不是孤立地处理定理,而是协调一个端到端的工作流程,涵盖规范说明、证明规划、逻辑审查、Lean证明构建、知识整理和验证,并支持独立监督和人工指导。其核心组成部分是基于证据的语义审计,用于评估形式化陈述是否忠实地保留了其非形式化规范。一个语言模型在局部对应关系、遗漏、范围和逻辑关系上构建结构化证据,而一个确定性验证器检查这些证据并产生可复现的判断。当接受一个实质性但可允许的偏差时,FYAN要求一个明确的证明转移义务,将形式化陈述与面向来源的解释联系起来。在每一阶段使用相同的模型(DeepSeek-V4.1-Flash),FYAN在严格的Lean检查下证明了143个FormalTCS定理中的86个,而通用智能体框架为69个,并将自然语言证明得分从0.501提高到0.851。在ConsistencyCheck上,其语义审计比直接LLM评判者捕捉到更多不一致的陈述,无论是在针对来源验证的标签上(召回率0.777对0.636)还是在原始标签上(0.873对0.820),并将其报告的每个不匹配定位到特定的假设、结论或范围。FYAN还构建了ODENumLib,一个用于常微分方程数值分析的9,355行Lean库。
英文摘要
We present FYAN, a human--AI harness for document-level mathematical formalization. Rather than treating theorems in isolation, FYAN coordinates an end-to-end workflow spanning specification, proof planning, logical review, Lean proof construction, knowledge curation, and validation, with support for independent supervision and human guidance. A central component is evidence-grounded semantic auditing, which assesses whether formal statements faithfully preserve their informal specifications. A language model constructs structured evidence over local correspondences, omissions, scope, and logical relations, while a deterministic validator checks this evidence and produces reproducible judgments. When a substantive but admissible deviation is accepted, FYAN requires an explicit proof-transfer obligation connecting the formal statement back to a source-facing interpretation. With the same model (DeepSeek-V4.1-Flash) in every stage, FYAN proves 86 of 143 FormalTCS theorems under a strict Lean check, against 69 for a general agent harness, and raises the natural-language proof score from 0.501 to 0.851. On ConsistencyCheck, its semantic audit catches more inconsistent statements than a direct LLM judge, both on labels verified against the source (recall 0.777 vs. 0.636) and on the original labels (0.873 vs. 0.820), and localizes each mismatch it reports to a specific hypothesis, conclusion, or scope. FYAN also built ODENumLib, a 9,355-line Lean library for the numerical analysis of ordinary differential equation.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。