发表机构
Computer Science Institute, Charles University; Institute of Formal and Applied Linguistics, Charles University; Department of Applied Mathematics, Charles University; Institute of Mathematics, Czech Academy of Sciences(查尔斯大学计算机科学研究所; 查尔斯大学形式与应用语言学研究所; 查尔斯大学应用数学系; 捷克科学院数学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文介绍多智能体开源系统Bolzano,通过并行证明智能体和验证智能体,在无人工指导下解决了约3800个开放问题中的约200个,并在STOC 2026论文中确认回答了四个问题。
AI 中文摘要
大型语言模型正越来越多地应用于数学研究,而此类研究的进展往往依赖于高效的证明搜索、增量式改进和细致的验证。我们描述了Bolzano,一个多智能体开源系统,它使用并行的证明智能体配合一个验证智能体,并维护一个人类可读的研究状态。在专家选定问题上的初步手动使用产生了8个结果,其证明已由领域专家检查。受这些案例研究的启发,我们在从四组论文中提取的约3,800个开放问题上运行了Bolzano,无需针对特定问题的人类指导,解决了约200个开放问题。其中一项实验使用了被理论计算机科学顶级会议STOC 2026录用的论文。在那里,我们回答了论文中提出的四个问题,并得到了作者们的确认。
英文摘要
Large language models are increasingly contributing to mathematical research, where progress often depends on efficient proof search, incremental improvements and careful verification. We describe Bolzano, a multi-agent open-source system that uses parallel prover agents with a verifier agent and maintains a human-readable research state. Initial manual use on expert-selected problems yielded 8 results whose proofs were checked by domain experts. Motivated by these case studies, we ran Bolzano without problem-specific human guidance on about 3,800 open problems extracted from four sets of papers, solving about 200 open problems. One experiment used papers accepted to STOC 2026, a top conference in theoretical computer science. There, we answered four questions raised in the papers, as confirmed by their authors.
CommentsAccepted at the 6th Workshop on Mathematical Reasoning and AI (MATH-AI), NeurIPS 2026