arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14582cs.AI

MathCoPilot:一种用于数学研究的人机共生范式的交互式系统

MathCoPilot: An Interactive System for Human-AI Symbiotic Paradigm of Mathematical Research

发表机构中国科学技术大学 · 中国科学技术大学数学科学学院
查看机构详情
  • University of Science and Technology of China(中国科学技术大学)
  • School of Mathematical Sciences, University of Science and Technology of China(中国科学技术大学数学科学学院)

机构由 AI 辅助整理,请以论文原文为准。

Junjie Zhang, Jiayu Liu, Wenbin Liu, Zhenya Huang, Doudou Wang, Yan Jiang, Leiye Xu, Tao Xiong, Wen Huang, Qi Liu, Guoping Hu, Enhong Chen, Mengping Zhang, Xiangdong Ye

首次发表
浏览论文内容

中文总结 AI 辅助

研究提出MathCoPilot这人机共生的数学研究系统,统一三项核心能力。通过它在特定子集和定理上比较四个大语言模型,发现当前模型处理本科问题成功率高,但在特定领域定理上仍面临挑战。

中文摘要 AI 辅助

现有的基于大语言模型的定理证明器在形式数学基准测试中取得了令人瞩目的成果,但它们仍局限于作为自主代理来证明给定命题。本文提出了MathCoPilot,这是一个人在回路系统,体现了一种新的数学研究人机共生范式,数学家掌控高层次数学方向,人工智能代理在持续的人类指导下进行详细的形式化和证明工作。MathCoPilot统一了三项核心能力:交互式工作台、自动证明技能编排、主题驱动的论文检索和自动形式化。使用MathCoPilot,在FormalMATH子集和两个需要深厚领域专业知识的实际偏微分方程定理上,系统地比较了四个最先进的大语言模型,评估它们产生经过验证的Lean 4证明和识别故意错误证明中的错误的能力。结果表明,虽然当前模型在有利的自动形式化条件下能以高成功率处理本科水平问题,但对于需要真正数学理解的特定领域定理仍存在重大挑战。

英文摘要

Existing LLM-based theorem provers have achieved impressive results on formal mathematics benchmarks, yet they remain confined to acting as autonomous agents that prove a stated proposition. In this paper, we propose MathCoPilot, a human-in-the-loop system that embodies a new human--AI symbiotic paradigm for mathematical research, in which the mathematician steers the high-level mathematical direction while AI agents carry out the detailed formalization and proof work under continuous human guidance. MathCoPilot unifies three core capabilities: (1) an interactive workbench where the mathematician and AI agents collaborate through a living proof blueprint that decomposes a proof into navigable steps the human can directly inspect, direct, and refine; (2) automated proving skill orchestration with adaptive knowledge base search and Lean-integrated iterative verification; and (3) topic-driven paper retrieval and automated formalization into a verified Lean knowledge base. Using MathCoPilot, we systematically compare four state-of-the-art LLMs, including Gemini~3.1~Pro, GPT-5.4, and Claude~Opus~4.7, on a FormalMATH subset and on two real PDE theorems requiring deep domain expertise, evaluating their ability to produce verified Lean~4 proofs and to identify errors in deliberately incorrect proofs. Our results show that while current models can handle undergraduate-level problems with high success rates under favorable autoformalization conditions, substantial challenges remain for domain-specific theorems requiring genuine mathematical understanding.

↑