发表机构
Singapore Management University; Monash University Malaysia; Daegu Gyeongbuk Institute of Science and Technology (DGIST); The Chinese University of Hong Kong; North Carolina State University; City University of Hong Kong; The Hong Kong University of Science and Technology; The University of Hong Kong(新加坡管理大学; 马来西亚莫纳什大学; 大邱庆北科学技术院; 香港中文大学; 北卡罗来纳州立大学; 香港城市大学; 香港科技大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出SWE-RPG编码智能体基准,评估发现主流编码智能体在该基准上平均解决率仅31.5%,隐性需求恢复是主要瓶颈,为改进编码智能体提供了方向。
AI 中文摘要
基于大语言模型的编码智能体正越来越多地被用于修改现有代码仓库,例如添加功能或修复漏洞。然而现有的仓库级基准通常仅评估最终补丁是否通过测试。满足用户请求需要一长串相互关联的推理和决策:智能体必须明确显性和隐性需求、制定基于仓库的实现计划,并将其转换为正确的代码。通过/不通过的结果无法描述失败的轨迹与正确补丁所需的需求和实现过程之间的偏差。为解决这一差距,我们推出SWE-RPG,这是一个仓库级基准,将可执行补丁评估与需求澄清和实现规划的验证基准真值(GT)相结合。这些中间基准真值支持对编码智能体在澄清、规划、代码生成和工件提交全过程的轨迹进行与基准真值对齐的回顾性诊断。SWE-RPG包含来自31个Python和Java仓库的163个任务,其中包括113个漏洞修复和50个功能添加。我们评估了3个编码智能体,包括Claude Code、Codex和OpenCode,以及6个大语言模型后端,包括Claude-Sonnet-5和GPT-5.6-Terra。结果显示,评估的主流编码智能体在现有仓库中实现用户请求仍然存在困难,在SWE-RPG上的平均解决率仅为31.5%。中间基准真值诊断进一步发现,隐性需求恢复是主要瓶颈,占智能体运行的24.5%至46.0%。这一结果表明,隐性需求恢复是改进编码智能体的关键候选方向。基准数据和评估代码可在此https URL获取。
英文摘要
Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs. Yet existing repository-level benchmarks typically evaluate only whether the final patch passes tests. Satisfying a user request requires a long chain of interdependent reasoning and decisions: an agent must recover explicit and implicit requirements, formulate a repository-grounded implementation plan, and translate it into correct code. A pass/fail outcome cannot characterize how an unsuccessful trajectory diverges from the requirements and implementation process needed for a correct patch. To address this gap, we introduce SWE-RPG, a repository-level benchmark that combines executable patch evaluation with validated ground-truth references (GTs) for (1) Requirement Clarification and (2) Implementation Planning. These intermediate GTs support retrospective, GT-aligned diagnosis of complete coding-agent trajectories across clarification, planning, code generation, and artifact submission. SWE-RPG comprises 163 tasks from 31 Python and Java repositories, including 113 bug fixes and 50 feature additions. We evaluate 3 coding agents, including Claude Code, Codex, and OpenCode, with 6 large language model backends, including Claude-Sonnet-5 and GPT-5.6-Terra. Results show that the evaluated popular coding agents still struggle to implement user requests in existing repositories, achieving an average resolved rate of only 31.5% on SWE-RPG. Intermediate-GT diagnosis further identifies implicit requirement recovery as the main bottleneck, accounting for 24.5%--46.0% of agent runs. This result suggests implicit-requirement recovery as a key candidate direction for improving coding agents. The benchmark data and evaluation code are available at https://github.com/Xin-Zhou-smu/SWE-RPG-Bench.
Comments9 pages