arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从零开始多智能体编码中作为一等公民的协作模式的实证研究

An Empirical Study of Coordination Mode as the First-Class Citizen in From-Scratch Multi-Agent Coding

Yanyu Ren, Yunfeng Bai, Xizheng Wang, Li Chen, Dan Li

arXiv 2607.27877首次发表:更新:

AI 中文总结

本研究构建多智能体编码基准MSEval,发现组织拓扑与模型能力对速度-成本-质量权衡影响相当,结构化流水线表现最优,为多智能体软件开发评估提供严格标准。

AI 中文摘要

多智能体氛围编码有望加速软件开发,但现有基准依赖合成环境,忽略实际时间与金钱成本,将推理与通信混为一谈,仅奖励表面完成度。我们引入多智能体从零开始评估基准MSEval,用于评估多智能体在真实任务上的编码性能。MSEval基于10个跨10个领域的真实全栈项目,采用分层需求与确定性评分规则评估性能;其执行引擎LegoGent测试10种协作拓扑,智能体通过定期同步间隔协调,并通过原生CI/CD流水线部署;自动评分器TAgent同步探测实现,共同衡量功能成功率、延迟及前缀缓存令牌成本。在100次运行中,MSEval显示组织拓扑与模型能力对速度-成本-质量权衡的影响相当;在相同任务与模型下,改变拓扑会使分数波动超30分, wall-clock时间翻倍;结构化流水线收敛最快且质量最高,而繁重的管理监督会降低性能。最终,MSEval确立了衡量多智能体团队实际构建软件方式的严格可复现标准,该基准已发布至指定网址。

英文摘要

Multi-agent vibe coding promises to accelerate software development, yet existing benchmarks rely on synthetic environments that ignore practical time and monetary costs, conflate reasoning with communication, and reward only superficial completion. We introduce multi-agent from-scratch evaluation benchmark, MSEval, evaluating multi-agent coding on real-world tasks. Grounded in 10 authentic, full-stack projects across 10 domains, MSEval scores performance using hierarchical requirements and deterministic rubrics. Its execution engine, LegoGent, tests 10 collaboration topologies where agents coordinate via periodic sync intervals and deploy through native CI/CD pipelines. Concurrently, the automated grader TAgent dynamically probes implementations to jointly measure functional success, latency, and prefix-cached token cost. Across 100 runs, MSEval reveals that organizational topology rivals model capability in shaping the speed--cost--quality trade-off. For identical tasks and models, varying the topology shifts scores by over 30 points and doubles wall-clock time. Structured pipelines converge fastest with the highest quality, whereas heavy managerial oversight degrades performance. Ultimately, MSEval establishes a rigorous, reproducible standard for measuring how multi-agent teams actually build software. The benchmark is released at https://github.com/robinren03/MSEval.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑