发表机构
Nanjing University; Central University of Finance and Economics; University of Illinois at Urbana-Champaign(南京大学; 中央财经大学; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过受控比较实验,考察了动态并发策略在长时程编码任务中的端到端效果,识别出13种失败模式和28种可观察模式,并揭示了其优势条件。
AI 中文摘要
随着编码代理从有限的软件工程任务向长时程开发推进,动态并发为扩展复杂开发任务提供了一种有前景的方式。在这种策略下,代理在执行过程中决定是否以及如何生成并发的子代理。模型能力在很大程度上决定了较短任务的结果,而长时程开发则使编排成为任务完成的核心。现有工作侧重于较短任务中的编码代理失败或预定义多代理工作流中的协作,对前沿代理在不同任务复杂度下的动态并发几乎没有提供见解。我们通过匹配的Codex、Claude Code和Kimi Code执行(启用或禁用该策略)的受控比较,将动态并发作为一种执行策略进行研究。在跨越多种任务复杂度和执行时域的354个任务和2124次执行中,我们评估了其端到端效果和调度行为,并分析了匹配的轨迹以刻画13种并发特定失败模式、28种可观察模式以及其提供优势的条件。
英文摘要
As coding agents advance from bounded software engineering tasks toward long horizon development, dynamic concurrency offers a promising way to scale complex development tasks. Under this policy, agents decide during execution whether and how to spawn concurrent sub-agents. Model capability largely determines outcomes on shorter tasks, whereas long horizon development makes orchestration central to task completion. Existing work, focused on coding agent failures on shorter tasks or collaboration in predefined multiagent workflows, offers little insight into dynamic concurrency in frontier agents across task complexity. We study dynamic concurrency as an execution policy through controlled comparisons of matched Codex, Claude Code, and Kimi Code executions with the policy enabled or disabled. Across 354 tasks and 2,124 executions spanning a range of task complexities and execution horizons, we evaluate its end to end effects and scheduling behavior, and analyze matched trajectories to characterize 13 concurrency specific failure modes, 28 observable patterns, and the conditions under which it provides an advantage.