对抗评审:基于结构化分歧的 grounded 智能体代码评审
Adversarial Review: Structured Disagreement for Grounded Agentic Code Review
- Cornell University(康奈尔大学)
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出对抗评审(AR)协议,仅用3个智能体实现结构化分歧的协作代码评审,在LiveCodeBench、SWE-PRBench等基准上优于多智能体基线,证明无需大量智能体即可完成高效代码评审。
AI中文摘要:
早期多智能体大语言模型系统常采用角色分离的团队架构,但在仓库级编码任务中,智能体数量的增加会导致收益递减。近期的替代方案将智能体视为被动工具(子智能体),这完全消除了智能体交互的优势。我们研究子智能体范式能否支持一种中间状态:最小化的智能体协作,而无需大型多智能体团队的开销。我们提出对抗评审(Adversarial Review, AR),一种最小化的协作代码评审协议,其中主编码智能体与评审智能体、批评智能体协同工作:评审智能体评估代码,批评智能体在主智能体编辑前通过结构化分歧对评审进行审核。在 LiveCodeBench 上,AR 在所有测试方法中达到最高通过率,仅使用 3 个智能体,优于 5 个智能体的基线。在 SWE-PRBench 上,朴素 AR 存在虚假共识失效模式,即智能体在缺乏充分证据的情况下达成一致,但通过一次添加明确分歧的提示迭代,AR 在所有测试方法中达到最高 F1 值。在 SWE-bench Verified 上,AR 在仓库级编码任务中也表现出优于基线的改进。综上,AR 表明协作代码评审不需要大量智能体或复杂的通信结构,它要求分歧是最小化、结构化且基于证据的。
英文摘要:
Early multi-agent LLM systems often used role-separated teams, yet scaling agent count yields diminishing returns on repository-level coding tasks. Recent alternatives treat agents as passive tools (subagents), yet this removes the benefits of agent interaction entirely. We study whether a subagent paradigm can support a middle ground: minimal agentic cooperation without the overhead of large multi-agent teams. We introduce Adversarial Review (AR), a minimal cooperative code-review protocol in which a main coding agent works with a reviewer and a critic agent. The reviewer evaluates code, while the critic audits the review through structured disagreement before the main agent edits. On LiveCodeBench, AR achieves the highest pass rate among tested methods, outperforming a five-agent baseline while using only three agents. On SWE-PRBench, naive AR exposes a false-consensus failure mode, where agents converge on agreement without sufficient evidence, but a single prompt iteration that adds disagreement explicitly achieves the highest F1 among tested methods. On SWE-bench Verified, AR also shows improvements over the baselines on repository-level coding tasks. Together, AR demonstrates that cooperative code review does not require many agents or complex communication structures: it requires that disagreement be minimal, structured, and evidence-grounded.