arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04862cs.AIcs.LG

GitSwarm:去中心化复合推理

GitSwarm: Decentralized Compounding Inference

Vedant Shah, Ankur Samanta, Paras Dahal, Mikhail Plekhanov, Carole-Jean Wu, Scott Yih, Remi Munos, Rob Fergus, Jakob Foerster, Ruslan Salakhutdinov, Sanjeev Aro… 展开作者

Vedant Shah, Ankur Samanta, Paras Dahal, Mikhail Plekhanov, Carole-Jean Wu, Scott Yih, Remi Munos, Rob Fergus, Jakob Foerster, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Aaron Courville, Anirudh Goyal

首次发表
浏览论文内容

中文总结 AI 辅助

GitSwarm提出复合推理范式,通过Git仓库持久化中间工作,使同质代理协作推进长时程任务,在多个基准上超越基线,证明推理计算可跨剧集累积复用。

中文摘要 AI 辅助

长时程问题求解和科学研究需要计算在多次尝试中不断累积。部分解决方案、实验发现和失败的方法可以为后续工作提供信息,然而大多数推理时计算是围绕单个轨迹或候选方案组织的,而非一个持久的可复用工作体。我们将这种范式称为复合推理:组织推理时计算,使得中间工作得以持久化,并能被后续计算检查、扩展、组合或挑战。我们在GitSwarm中实例化了复合推理,这是一个异步系统,其中同质代理独立决定如何推进任务,同时通过结构化的持久记忆进行协作。代理在一个共享的、可分支的Git仓库中探索、实验、验证、改进和综合先前的工作。原子提交保留了中间产物,而显式的语义依赖记录了后续贡献如何跨分支构建于先前工作之上。我们在长时程问题求解和持续的GPU支持的实验研究上评估了GitSwarm。在IMOProofBench-Advanced上,GitSwarm使用GPT-5.5在一次运行中解决了全部30个问题。在ProgramBench上,它达到了79.4%的平均分数,而在所述预算下最强报告的基线为65.1%。在三个神经架构研究任务(Residual Matrix Transformer、Looped Transformer、NanoChat)中,GitSwarm通过连续实验改进了起始架构。除了最终性能,我们衡量了计算是否累积:在ProgramBench上,94.7%的贡献被后续构建,而所选解决方案的祖先覆盖了贡献图的82-93%。这些结果表明,推理时计算可以在原本独立的剧集之间累积,形成一个不断演进的工作体,供后续推理复用。

英文摘要

Long-horizon problem solving and scientific research require computation to accumulate across successive attempts. Partial solutions, experimental findings, and unsuccessful approaches can inform later work, yet most inference-time computation is organized around individual trajectories or candidates rather than a persistent body of reusable work. We call this paradigm compounding inference: organizing inference-time computation so that intermediate work persists and can be inspected, extended, combined, or challenged by subsequent computation. We instantiate compounding inference in GitSwarm, an asynchronous system where homogeneous agents independently decide how to advance a task while collaborating through structured persistent memory. Agents explore, experiment, verify, refine, and synthesize previous work in a shared, branch-able Git repository. Atomic commits preserve intermediate artifacts, while explicit semantic dependencies record how later contributions build on work across branches. We evaluate GitSwarm on long-horizon problem solving and sustained GPU-backed experimental research. On IMOProofBench-Advanced, GitSwarm solves all 30 problems in one run using GPT-5.5. On ProgramBench, it achieves a $79.4\%$ mean score, versus $65.1\%$ for the strongest reported baseline under the stated budget. On three neural architecture research tasks (Residual Matrix Transformer, Looped Transformer, NanoChat), GitSwarm improves upon the starting architectures through successive experimentation. Beyond final performance, we measure whether computation accumulates: on ProgramBench, $94.7\%$ of contributions are subsequently built upon, while the selected solution's ancestry covers $82-93\%$ of the contribution graph. These results show that inference-time computation can accumulate across otherwise independent episodes, forming an evolving body of work that subsequent inference can reuse.

发表机构

  • Meta Superintelligence Labs(Meta超级智能实验室)
  • Mila - Québec AI Institute(米拉-魁北克人工智能研究所)
  • Université de Montréal(蒙特利尔大学)
  • Columbia University(哥伦比亚大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑