发表机构
Baidu(百度)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出BulkPR-Bench基准,针对交互PR队列治理,实验显示现有模型在该任务上表现优于顺序基准,但仍存在较大提升空间。
AI 中文摘要
编码智能体基准日益覆盖长周期、端到端及交互式开发,但通常仅保留一个请求结果或固定变更序列。顺序策略可一次处理一个拉取请求(PR)队列候选,但当队列中的PR存在交互时,最大化安全交付需联合决定哪些变更应合并及合并顺序。我们推出BulkPR-Bench,这是一个可执行基准,要求智能体在滚动发布协议下恢复关键PR关系,并以可执行顺序返回一个大型安全子集。该套件包含18个真实仓库冻结快照上的581个新生成的候选PR。注册的逐状态仓库执行(包括隐藏的安全检查)验证黄金关系图;精确的神谕则计算最大安全子集。我们的主要指标是关系交付分数(RDS),用于评估从已实现的合并轨迹中关系组的安全交付和正确拒绝;全局安全门控产出(Global-SGY)则单独测量整个队列计划的严格交付。在批处理大小K=32的缓冲主协议下,六个模型中最高的三个RDS估计值分别为66.6%、62.0%和57.9%,而最强的顺序基准为53.1%。324次模型运行中仅有8次完全完成队列。关键关系召回率在35.2%至57.7%之间,且提供黄金关系的诊断运行显示仍有较大提升空间。因此,关系组上的增益尚未转化为可靠的整个队列治理能力。
英文摘要
Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcome or a fixed change sequence. Sequential policies can process a pull-request (PR) queue one candidate at a time, but when queued PRs interact, maximizing safe delivery can require jointly deciding which changes to merge and in what order. We introduce BulkPR-Bench, an executable benchmark in which an agent must recover consequential PR relations and return a large safe subset in executable order under a rolling-release protocol. The suite contains 581 newly authored candidate PRs on frozen snapshots of 18 real repositories. Registered state-by-state repository execution, including hidden safety checks, validates the gold relation graph; an exact oracle then computes the largest safe subset. Our primary metric, Relational Delivery Score (RDS), scores safe delivery and correct rejection over relation groups from the realized merge trace; Global Safety-Gated Yield (Global-SGY) separately measures strict delivery of the realized whole-queue plan. Under the buffered primary protocol with batch size $K=32$, the three highest RDS estimates among the six models are 66.6%, 62.0%, and 57.9%, compared with 53.1% for the strongest sequential baseline. Only 8 of 324 model runs complete a queue exactly. Critical-relation recall ranges from 35.2% to 57.7%, and diagnostic runs supplied with the gold relations show substantial remaining headroom. Gains on relation groups therefore do not yet translate into dependable whole-queue governance.
Comments12 pages, 5 figures. Artifact: https://github.com/Eureka246/BulkPR-Bench-Release ; archived artifact: https://doi.org/10.5281/zenodo.21717780