arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Stellar Colosseum:面向数学与理论计算机科学长时程研究的多智能体框架

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

Honghao Lin, David P. Woodruff, Yuan Deng, Jieming Mao, Song Zuo, Vahab Mirrokni

arXiv 2609.15983首次发表:更新:

发表机构

Google Research; Carnegie Mellon University(谷歌研究院; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长时程数学与理论计算机科学研究中语言模型不可靠的问题,提出多智能体框架 Stellar Colosseum,通过策略探索、就绪门控、并行候选生成与证伪聚合等方法,在多个基准上取得新结果,并集成到 Google Antigravity 框架中。

AI 中文摘要

语言模型能够生成看似合理的短证明,但在长时程研究问题上可能仍然不可靠,因为这类问题的进展依赖于一系列不确定且相互依赖的决策。我们提出 Stellar Colosseum,一个与模型无关的框架,用于在数学和理论计算机科学研究中分配推理资源。Colosseum 在证明构建之前探索替代策略,使用就绪门控(readiness gate)来决定一条路线何时成熟到可以分解,将证明计划表示为相互依赖的章节级子问题,并将验证器的发现路由回论证中受影响的部分。在这些阶段中,它并行生成候选方案,用有针对性的证伪(falsification)攻击它们,并通过重叠随机样本树聚合(overlapping random-sample tree aggregation)将候选方案及其批评意见合并为单一研究产物。Colosseum 工作流也已作为“长证明”(Long Proof)模式集成到 Google Antigravity 的 Teamwork 框架中。我们通过开放式研究以及在定理证明和竞争性编程基准上的评估来展示 Colosseum 的能力。使用 Colosseum 配合 Gemini 3.1 Pro,我们获得了若干新结果,解决了发表在 FOCS 和 JMLR 等顶级会议和期刊上的论文中提出的开放问题。在 TCS-Bench(一个从 FOCS、STOC 和 SODA 发表的论文中提取的研究级定理证明任务基准)上,Colosseum 使用 Gemini 3.1 Pro 和 Gemini 3.7 Flash 达到了 71.0% 的准确率。在另一次使用 Gemini 3.1 Pro 的 Codeforces 评估中,带执行反馈的面向证明的流水线解决了 222 个问题中的 218 个。

英文摘要

Language models can produce plausible short proofs, but may still be unreliable on long-horizon research problems, where progress depends on a sequence of uncertain and interdependent decisions. We introduce Stellar Colosseum, a model-agnostic harness for allocating inference across research in mathematics and theoretical computer science. Colosseum explores alternative strategies before proof construction, uses a readiness gate to decide when a route is mature enough to decompose, represents the proof plan as interdependent section-level subproblems, and routes verifier findings back to the affected part of the argument. Across these stages, it generates candidates in parallel, attacks them with targeted falsification, and combines candidates and their critiques into a single research artifact through overlapping random-sample tree aggregation. The Colosseum workflow has been integrated into Google Antigravity's Teamwork framework as the Long Proof pattern. We demonstrate the capabilities of Colosseum through open-ended research and evaluations on theorem-proving and competitive programming benchmarks. Using Colosseum with Gemini 3.1 Pro, we obtain several new results that address open problems arising from papers published at top venues such as FOCS and JMLR. On TCS-Bench, a benchmark of research-level theorem-proving tasks drawn from papers published at FOCS, STOC, and SODA, Colosseum achieves 71.0% accuracy using Gemini 3.1 Pro and Gemini 3.7 Flash. In a separate Codeforces evaluation using Gemini 3.1 Pro, the proof-oriented pipeline with execution feedback solves 218 of 222 problems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑