arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BulkPR-Bench:针对交互拉取请求的队列级治理基准

BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests

Zetong Xiong, Qiao Zhao, Jun Zhang, Xueying Lyu, Zhi Li, Yixiang Tu, Xiaowen Yang, Yunjie Zhang, Yufeng Wang, Zhe Zhang, Kaize Yu, Hanwen Du, Zhongkai Sun, Zhuoxin Liu, Zekun Lin, Jianwen Yang, Ruining Chen, Ying Zhang, Tingxuan Pan, Ke Chen, Shubin Han, Chuanhao Sun, Yehua Yang

arXiv 2608.02685首次发表:更新:

发表机构

Baidu(百度)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出BulkPR-Bench基准,针对交互PR队列治理,实验显示现有模型在该任务上表现优于顺序基准,但仍存在较大提升空间。

AI 中文摘要

编码智能体基准日益覆盖长周期、端到端及交互式开发,但通常仅保留一个请求结果或固定变更序列。顺序策略可一次处理一个拉取请求(PR)队列候选,但当队列中的PR存在交互时,最大化安全交付需联合决定哪些变更应合并及合并顺序。我们推出BulkPR-Bench,这是一个可执行基准,要求智能体在滚动发布协议下恢复关键PR关系,并以可执行顺序返回一个大型安全子集。该套件包含18个真实仓库冻结快照上的581个新生成的候选PR。注册的逐状态仓库执行(包括隐藏的安全检查)验证黄金关系图;精确的神谕则计算最大安全子集。我们的主要指标是关系交付分数(RDS),用于评估从已实现的合并轨迹中关系组的安全交付和正确拒绝;全局安全门控产出(Global-SGY)则单独测量整个队列计划的严格交付。在批处理大小K=32的缓冲主协议下,六个模型中最高的三个RDS估计值分别为66.6%、62.0%和57.9%,而最强的顺序基准为53.1%。324次模型运行中仅有8次完全完成队列。关键关系召回率在35.2%至57.7%之间,且提供黄金关系的诊断运行显示仍有较大提升空间。因此,关系组上的增益尚未转化为可靠的整个队列治理能力。

英文摘要

Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcome or a fixed change sequence. Sequential policies can process a pull-request (PR) queue one candidate at a time, but when queued PRs interact, maximizing safe delivery can require jointly deciding which changes to merge and in what order. We introduce BulkPR-Bench, an executable benchmark in which an agent must recover consequential PR relations and return a large safe subset in executable order under a rolling-release protocol. The suite contains 581 newly authored candidate PRs on frozen snapshots of 18 real repositories. Registered state-by-state repository execution, including hidden safety checks, validates the gold relation graph; an exact oracle then computes the largest safe subset. Our primary metric, Relational Delivery Score (RDS), scores safe delivery and correct rejection over relation groups from the realized merge trace; Global Safety-Gated Yield (Global-SGY) separately measures strict delivery of the realized whole-queue plan. Under the buffered primary protocol with batch size $K=32$, the three highest RDS estimates among the six models are 66.6%, 62.0%, and 57.9%, compared with 53.1% for the strongest sequential baseline. Only 8 of 324 model runs complete a queue exactly. Critical-relation recall ranges from 35.2% to 57.7%, and diagnostic runs supplied with the gold relations show substantial remaining headroom. Gains on relation groups therefore do not yet translate into dependable whole-queue governance.

Comments12 pages, 5 figures. Artifact: https://github.com/Eureka246/BulkPR-Bench-Release ; archived artifact: https://doi.org/10.5281/zenodo.21717780

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑