arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgentRadio:面向长 horizon 多智能体协作的被动感知机制

AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration

Xinxing Ren, Qianbo Zang, Ziyan Wang, Caelum Forder, Suman Deb, Peter Carroll, Zekun Guo

arXiv 2607.28430首次发表:更新:

AI 中文总结

针对现有多智能体系统仅能在阶段间交换信息的局限,提出异步消息传递层 AgentRadio,使编码智能体保持被动感知,在 SWE-Atlas QnA 基准上多智能体任务解决率显著优于单个智能体及 Opus 4.8 版 Claude Code。

AI 中文摘要

理解大型代码库是大型语言模型(LLM)智能体的长 horizon 任务,回答单个问题可能需要构建并运行软件、跨文件追踪执行过程,以及在数十分钟内整合证据。在针对生产仓库的长 horizon 问题基准 SWE-Atlas QnA 上,单个 Claude Code 智能体(Opus 4.6)仅能解决 32.3% 的任务。将工作分配给拥有干净上下文的智能体可缓解这一局限,但代码理解的子任务相互依赖,一个智能体的发现可能会改写另一个智能体的任务,因此智能体必须在执行过程中而非仅在阶段边界进行协调。现有多智能体系统仅支持在阶段间通过阶段性交接或同步轮次实现此类信息交换,通信与工作仍互斥,执行中途的发现需等到下一个边界才能共享。我们提出 AgentRadio,这是一个异步消息传递层,为编码智能体框架配备三个原语:线程、消息和等待提及。其中等待提及作为后台任务运行,可在不中断前台工作的情况下呈现队友的消息,使每个智能体保持被动感知,将新发现整合到正在进行的任务中。在分工与协商的五阶段协议下,由 AgentRadio 组织的四个智能体解决了 62.1% 的任务,比单个智能体高出 29.8 个百分点,也优于采用更新 Opus 4.8 的 Claude Code(57.2%)。评分规则级分析显示,性能增益随任务难度增大而提升,与中途修正作为底层机制的情况一致。我们的代码可在 https URL 获取。

英文摘要

Understanding large codebases is a long-horizon task for Large Language Model (LLM) agents: answering a single question can require building and running the software, tracing execution across files, and synthesizing evidence over tens of minutes. On SWE-Atlas QnA, a benchmark of long-horizon questions over production repositories, a single Claude Code agent (Opus 4.6) resolves only 32.3% of tasks. Dividing the work among agents with clean contexts mitigates this limitation. However, the subtasks of code comprehension are interdependent. One agent's findings can rewrite another's task, so agents must coordinate during execution, not only at phase boundaries. Existing multi-agent systems support such exchange only between phases, through staged handoffs or synchronized rounds. Communication and work remain mutually exclusive. A discovery made mid-execution cannot be shared until the next boundary. We present AgentRadio, an asynchronous message-passing layer that equips coding-agent harnesses with three primitives: threads, messages, and waiting for mentions. The last runs as a background task, surfacing teammates' messages without interrupting foreground work, so each agent remains passively aware of its peers and folds new findings into its ongoing task. Under a five-phase protocol of division of labor and negotiation, four agents organized by AgentRadio resolve 62.1% of tasks, 29.8 points above a single agent and above Claude Code with the newer Opus 4.8 (57.2%). Rubric-level analysis shows the gain growing with task difficulty, consistent with mid-course correction as the underlying mechanism. Our code is available at https://github.com/Coral-Protocol/AgentRadio.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑