发表机构
University of Houston; New York University; Texas A&M University; UCSD; Worcester Polytechnic Institute; Oregon State University(休斯顿大学; 纽约大学; 德克萨斯农工大学; 加州大学圣地亚哥分校; 伍斯特理工学院; 俄勒冈州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出依赖为中心的异步多智能体协作基准AsynCodeBench,用ADPR和DRS度量协作,发现编码能力与协作能力存在差距,并识别出跳跃窗口模式。
AI 中文摘要
多智能体编程已成为软件工程中一个日益活跃的方向,其中复杂的开发任务被分解到多个专门处理问题不同部分的智能体上。尽管从个体问题求解转向分布式协作,多智能体系统仍然缺乏对协作的直接度量,并且主要通过继承自单智能体编程的任务级结果来评估,这混淆了个体编码能力与跨智能体协调能力。我们引入了AsynCodeBench,一个面向异步多智能体软件工程的以依赖为中心的基准测试,它用显式的依赖图和可执行的依赖检查器来表示每个任务。通过这种依赖追踪过程,我们提出了两个互补的度量:异步依赖通过率(ADPR),它衡量最终满足的跨智能体依赖的数量;以及依赖解析步骤(DRS),它衡量每个依赖在执行过程中首次被满足的时间。AsynCodeBench包含来自真实世界代码库的19个任务,暴露了52个有向依赖作为评估跨智能体协作的显式单元。跨模型家族、规模和世代的实验揭示了编码能力与协作能力之间的明显差距:编码性能的提升并不一定转化为更强的协作,任务级指标可能与依赖级协作度量产生显著分歧。依赖轨迹分析进一步表明,成功的协调往往不是逐渐出现的,而是通过集中的爆发,在执行的短时间片段内许多依赖被解析,我们将这种模式称为跳跃窗口。
英文摘要
Multi-agent coding has emerged as an increasingly active direction in software engineering, where complex development tasks are decomposed across multiple specialized agents working on different parts of the problem. Despite the shift from individual problem solving to distributed collaboration, multi-agent systems still lack a direct measure of collaboration and are largely evaluated through task-level outcomes inherited from single-agent coding, conflating individual coding capability with cross-agent coordination. We introduce AsynCodeBench, a dependency-centric benchmark for asynchronous multi-agent software engineering that represents each task with an explicit dependency graph and executable Dependency Checkers. Through this dependency-tracking process, we propose two complementary measures: Asynchronous Dependency Pass Rate (ADPR), which measures how many cross-agent dependencies are ultimately satisfied, and Dependency Resolution Step (DRS), which measures when each dependency first becomes satisfied during execution. AsynCodeBench comprises 19 tasks from real-world repositories, exposing 52 directed dependencies as explicit units for evaluating cross-agent collaboration. Experiments across model families, scales, and generations reveal a clear gap between coding and collaboration capability: improvements in coding performance do not necessarily translate into stronger collaboration, and task-level metrics can diverge substantially from dependency-level collaboration measures. Dependency-trajectory analysis further reveals that successful coordination often emerges not gradually, but through concentrated bursts in which many dependencies become resolved over a short portion of the execution trajectory, a pattern we term a hopping window.