arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单独通过,合并失败:并行LLM智能体开发中的语义协调基准测试

Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development

Haocheng Xia, Eugene Wu, Yongjoo Park

arXiv 2609.25396首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Columbia University(伊利诺伊大学厄巴纳-香槟分校; 哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出stale基准,评估并行编码智能体合并时的语义协调失败,发现真实拉取请求干扰少,而构造任务干扰率高,消息可恢复多数失败。

AI 中文摘要

并行编码智能体可以生成单独工作时正常但合并时失败的补丁。这种情况发生在一个智能体更改了另一个智能体仍然依赖的接口或规则时。我们使用stale(一个语义协调基准)来研究这些失败。我们的评估对每个补丁单独及其组合运行相同的测试,仅统计由补丁组合引入的失败。我们使用三个层级:具有受控接口变化的合成任务、合并的拉取请求对,以及使用真实Django辅助函数的构造任务。在834次对417个挖掘的Django对的运行中,在纠正评分程序后只有一次显示出干扰。在使用12个Django辅助函数的构造任务中,干扰发生在97%的运行中。一条描述已完成的并发更改的消息恢复了82%的运行。审查过的拉取请求可能包含很少的未解决的并行更改,即使智能体在使用真实代码的受控任务上失败。构造的失败率并不能估计这些问题在实践中发生的频率。

英文摘要

Parallel coding agents can produce patches that work alone but fail when merged. This happens when one agent changes an interface or rule that another agent still relies on. We study these failures with stale, a benchmark for semantic coordination. Our evaluation runs the same tests on each patch alone and on their combination, counting only failures introduced by combining the patches. We use three tiers: synthetic tasks with controlled interface changes, pairs of merged pull requests, and constructed tasks that use real Django helpers. Among 834 runs on 417 mined Django pairs, only one showed interference after correcting the grading procedure. On constructed tasks using 12 Django helpers, interference occurred in 97% of runs. A message describing the completed concurrent change recovered 82% of runs. Reviewed pull requests may contain few unresolved parallel changes, even when agents fail on controlled tasks using real code. The constructed failure rates do not estimate how often these problems occur in practice.

Comments6 pages, accepted to The 2nd Workshop on Explainable and Reliable Software Systems (EXPRESS 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑