arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28281cs.AI

LoopArena:将模型作为循环工程运行时控制器的基准测试

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering

Yi Wang, Haopeng Zhang, Chengxiang Huang, Rui Dai, Kaikui Liu, Piotr Koniusz, Xiangxiang Chu

首次发表
浏览论文内容

中文总结 AI 辅助

研究推出LoopArena基准,评估模型作为控制器指导编码智能体完成长期任务的能力,发现完整任务上最佳严格成功率为24.69%,推理成本平均降低64.4%。

中文摘要 AI 辅助

循环工程是围绕编码智能体组织开发工作的新兴实践。从业者不再手动编写每个提示,而是设计循环来监控进度、分配工作、运行检查,并决定智能体下一步应执行的操作。即使拥有强大的编码智能体,循环也可能信任过时的进度记录、跳过必要的验证、将预算花在错误方向,或在任务安全提交前停止。然而,一次端到端运行的最终结果无法判断成功或失败是反映了循环的指导还是编码智能体执行任务的能力。我们推出LoopArena,这是一个用于评估模型指导独立编码智能体完成长期任务能力的基准测试。被评估的模型是“控制器”:在每次编码轮次后,它会收到运行的结构化摘要,并指示独立的、固定的编码智能体“工作者”下一步要做什么或验证什么,或决定是否停止。LoopArena在三个执行范围和成本不同的互补设置中评估此能力:I型通过无需在评估时运行工作者的、经执行验证的问题对下一步循环契约选择进行评分;II型对完整任务的选定片段执行重复控制;III型从原始状态评估配对的完整任务。在完整任务上,观察到的最佳严格成功率为24.69%,表明长期循环控制仍有很大改进空间。在各控制器中,估计推理成本的配对降低平均为64.4%,且II型在核心标准下产生相似的排序(斯皮尔曼ρ=0.9747)。我们在该httpsURL发布基准数据和评估代码。

英文摘要

Loop Engineering is emerging as a practice for organizing development work around coding agents. Instead of writing each prompt by hand, practitioners design loops that monitor progress, assign work, run checks, and decide what the agent should do next. Even with a capable coding agent, a loop may trust a stale progress note, skip needed verification, spend its budget in the wrong direction, or stop before the task is safe to submit. Yet the final outcome of one end-to-end run cannot tell whether success or failure reflects the loop's guidance or the coding agent's ability to carry out the task. We introduce LoopArena, a benchmark for evaluating how well one model can guide a separate coding agent through a long-running task. The model under evaluation is the \textbf{Controller}: after each coding round, it receives a structured summary of the run and instructs a separate, fixed coding agent, the \textbf{Worker}, on what to do or verify next, or decides whether to stop. LoopArena evaluates this ability in three complementary settings that differ in execution scope and cost. Type I scores next-step Loop Contract selection through execution-validated questions without running the Worker at evaluation time. Type II executes repeated control over a selected slice of a full task, while Type III evaluates the paired full task from its original state. On full tasks, the best observed Strict Success Rate is \textbf{24.69\%}, leaving substantial room for improvement in long-horizon loop control. Across Controllers, the paired reduction in estimated inference cost averages \textbf{64.4\%}, and Type II produces a similar ordering under the main Core criterion (Spearman's \(ρ=\textbf{0.9747}\)). We release the benchmark data and evaluation code at https://github.com/AMAP-ML/LoopArena .

发表机构

  • DreamX Team, Alibaba Group(达摩院团队,阿里巴巴集团)
  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • UNSW Sydney(新南威尔士大学悉尼分校)
  • Data61, CSIRO(联邦科学与工业研究组织数据61研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑