arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07261cs.AI

验证并行编码智能体的协调性:NP-Bench 与调度规划器

Verifying Coordination in Parallel Coding Agents: NP-Bench and a Scheduling Planner

Sumanyu Muku

首次发表
浏览论文内容

中文总结 AI 辅助

针对并行编码智能体协调失败问题,提出基于调度规划的主动协调方法,将工作划分和合并排序前置,在NP-Bench基准上将干净集成率从1/9提升至9/9,并消除合并冲突。

中文摘要 AI 辅助

一个编码智能体团队在逐个检查时可能看起来没有问题,但作为一个团队却会失败:每个智能体都通过了自己的测试,而合并后的结果却是损坏的,单智能体评估永远无法发现这一问题。当团队在一个代码库上并行运行多个 LLM 编码智能体时,智能体之间会发生冲突:两个智能体重写了同一个函数,一个智能体基于队友刚刚更改的合约进行编码,集成在工作完成后失败。大多数协调工具是反应式的(监视冲突,然后发出警告),但在智能体的速度下,警告总是在浪费的编辑之后才到达。我们将问题重新定义为调度问题:获取每个工作项的声明范围,将工作划分为不相交的范围,并沿着生产者->消费者图对合并进行排序,所有这些都在事前完成。我们将此规划器构建到 Nerveplane 中,并使用 NP-Bench 对其进行评估,NP-Bench 是一个基于环境的三臂基准(无协调;反应式检测;主动规划),它在确定性模拟和实时智能体两种情况下,都通过真实的 git 合并来验证集成。该规划器将干净集成从 1/9 场景提升到 9/9 场景,并将合并冲突从 13 减少到 0,且这一优势随智能体数量的增加而扩大。在一个实时的破坏性合约变更中,它在所有种子下都挽救了两个基线都失败的结果:干净集成率从前沿模型的 0(无协调和反应式检测)提升到 1.0,小型模型提升到 0.6,同时智能体尊重分配的范围(0/5 泄漏)。跨会话记忆将强模型和弱模型的重复错误率从 1.00 降至 0.00。我们还报告了一个负面结果:在窗口适配规模下,将事实路由给智能体并不能挽救长上下文准确性;其价值在于成本和容量,而非注意力。在两个能力层级和两个供应商中,随着模型变得更强,这一收益并未缩小,因为它来自工作的分配方式,而非模型的推理能力。

英文摘要

A team of coding agents can look fine agent by agent yet fail as a team: each passes its own tests while the merged result is broken, and single-agent evaluation never catches it. As teams run several LLM coding agents in parallel on one codebase, the agents collide: two rewrite the same function, one codes against a contract a teammate just changed, and integration fails after the work is done. Most coordination tools react (watch for a conflict, then warn), but at agent speed the warning arrives after the wasted edit. We recast the problem as scheduling: take each work item's declared scope, partition the work into disjoint scopes, and order merges along the producer->consumer graph, all up front. We build this planner into Nerveplane and evaluate it with NP-Bench, an environment-grounded three-arm benchmark (no coordination; reactive detection; proactive planning) that verifies integration off a real git merge, both in a deterministic simulation and with live agents. The planner lifts clean-integration from 1/9 to 9/9 scenarios and cuts merge conflicts from 13 to 0, with a gap that grows in the number of agents. On a live breaking contract change it rescues an outcome both baselines miss on every seed: the clean-integration rate rises from 0 (no coordination and reactive detection) to 1.0 on a frontier model and 0.6 on a small one, while agents respect assigned scopes (0/5 leakage). A cross-session memory drops the repeated-mistake rate from 1.00 to 0.00 on strong and weak models alike. We also report a negative result: routing facts to agents does not rescue long-context accuracy at window-fitting scales; its value is cost and capacity, not attention. Across two capability tiers and two vendors, the benefit did not shrink as models got stronger, because it comes from how work is allocated, not model reasoning.

发表机构

  • Amira Learning

机构由 AI 辅助整理,请以论文原文为准。

↑