arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38345cs.SEcs.AIcs.CLcs.LG

OpenCollab:一个具有可编程协作与可控运行时的多智能体编码框架

OpenCollab: A Multi-Agent Coding Framework with Programmable Collaboration and Controllable Runtime

Chun-Wah Hsu, Kai Gong, Yu Wu, Xianhe Chen, Mengyang Liu, Jie Li, Hanyu Li, Zhixuan Liu, Naisheng Tang, Jiaying Chi, Ziheng Fan, Xuning He, Xiaokang Yang, Xue Jiang, Yihong Dong

首次发表
浏览论文内容

中文总结 AI 辅助

OpenCollab提出统一多智能体编码框架,通过可编程协作与可控运行时量化组织遵循度,双编码器工作流实现新SOTA性能,并节省令牌。

中文摘要 AI 辅助

多智能体编码系统旨在通过协作解决复杂的软件工程任务。然而,现有的评估通常假设配置的组织结构会被忠实遵循,而现实情况并非如此。这种行为上的差距,加上底层系统组件的差异,阻碍了对所观察到的性能提升进行清晰归因。为此,我们提出了OpenCollab,一个多智能体编码框架,它为可编程协作和可控运行时提供了统一的基础设施。具体而言,OpenCollab统一了组织设计,在共享运行时上强制执行实验控制,并通过细粒度的事件流跟踪执行过程。在此基础上,我们定义了“遵循度”(Adherence)来量化所声明的组织是否真正得以实现。我们的实验表明,智能体在不同配置下的协作方式差异显著:改变任何单一维度都会使遵循度从47.2%变化到高达97.2%。此外,广泛的智能体编码基准测试表明,基于OpenCollab构建的双编码器工作流相较于主流工具(如Mini-SWE-agent、Codex CLI和Claude Code)确立了新的最优性能(SOTA),这表明精心设计的组织可以超越强大的现有工具,而OpenCollab的单智能体配置在所有评估套件中使用的令牌数最少。OpenCollab建立了一个统一的多智能体基础设施,便于可编程协作和受控的因果评估。

英文摘要

Multi-agent coding systems are designed to tackle complex software engineering tasks through collaboration. However, existing evaluations typically assume configured organizations are followed faithfully, whereas reality differs. This behavioral gap, combined with differences in underlying system components, prevents clear attribution of observed gains. To this end, we introduce OpenCollab, a multi-agent coding framework that provides a unified infrastructure for programmable collaboration and controllable runtime. Specifically, OpenCollab unifies organization design, enforces experimental control on a shared runtime, and tracks execution through fine-grained event streams. On this basis, we define Adherence to quantify whether the declared organization is actually realized. Our experiments reveal that agents collaborate very differently across configurations: changing any single dimension shifts Adherence, from 47.2% to as high as 97.2%. Furthermore, extensive agentic coding benchmarks show that a two-coder workflow built on OpenCollab establishes new SOTA performance compared to the mainstream harnesses such as Mini-SWE-agent, Codex CLI, and Claude Code, showing that a well-designed organization can outperform strong existing harnesses, while OpenCollab's single-agent configuration uses the fewest tokens across all evaluated suites. OpenCollab establishes a unified multi-agent infrastructure for easy programmable collaboration and controlled causal evaluation.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑