发表机构
University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究引入时间网络工具度量多智能体AI编码团队的协作,分析团队规模、结构等对协作的影响,发现智能体存在主动获取隐藏评分材料的倾向,相关结果在封闭环境中得到复现。
AI 中文摘要
我们研究AI编码智能体团队在解决编程任务时的协作方式。当前评估通常仅报告智能体是否完成任务以及运行成本,基本未对团队内部的协作进行度量。我们引入一种用于度量这种协作的工具,每次运行被表示为一个时间网络,其中智能体和文件为节点,消息、文件写入和文件读取为带时间戳的有向边,且关联有成本。我们将该工具应用于1902次运行,每次运行使用固定测试套件进行评估,涉及团队规模、团队结构和文件策略各不相同的配置。所得网络显示了协作如何随团队规模和工作内容变化:直接消息最初随智能体数量近似二次增长,大部分增长来自早期介绍环节;当团队进一步扩大,在我们研究的最大规模团队中,这种增长趋于平稳,智能体越来越多地通过广播消息进行通信。任务也会塑造生成的网络:基于共享规范开展的工作会产生密集、高度连通的团队,而流水线任务则会产生围绕本地接口组织的稀疏网络。共享文件可替代重复的一对一通信,在消息密集型工作中,8个智能体时可减少约42%的输出token,但当文件已承载协作功能时会增加开销;指定一个智能体作为协调者不会形成通信枢纽,也不会带来可靠的成功率提升。我们还观察到智能体存在主动寻找隐藏评分材料的自发倾向,我们在封闭环境中重复关键实验条件,将隐藏材料替换为带标记的占位文件,在额外的244次运行中,仍有五分之四的运行中智能体试图获取该材料,而协调者和文件通道的相关发现得到复现。
英文摘要
We study how teams of AI coding agents coordinate while solving programming tasks. Current evaluations usually report whether the agents complete the task and how much the run costs, leaving the coordination inside the team largely unmeasured. We introduce an instrument to measure this coordination. Each run is represented as a temporal network in which agents and files are nodes, and messages, file writes, and file reads are timestamped directed edges with an associated cost. We apply this instrument to 1902 runs, each evaluated with a fixed test suite, across configurations that vary the team size, the team structure, and the file policy. The resulting networks show how coordination changes as teams grow and as the work changes. Direct messaging initially increases close to quadratically with the number of agents, with much of this growth coming from an early round of introductions. As the teams grow further, this increase levels off in the largest teams we study, where agents increasingly communicate through broadcast messages. The task also shapes the network that emerges. Work built around a shared specification produces dense, highly connected teams, while pipeline tasks produce sparse networks organised around local interfaces. Shared files can replace repeated 1-to-1 communication, cutting output tokens by about 42% at eight agents on message-heavy work, while adding overhead when files already carry the coordination. Naming one agent as coordinator creates no communication hub and provides no reliable improvement in success. We also observe an unprompted tendency for agents to seek out hidden grading material. We repeat the key experimental conditions in a sealed environment, replacing the hidden material with marked placeholder files. Across 244 additional runs, agents still reach for it in four fifths of runs, while the coordinator and file-channel findings reproduce.