arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00267cs.SEcs.CL

LoopsBench:在基准测试编码智能体中从 harness 工程转向 loop 工程

LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation

Han Li, Zhemin Fang, Rili Feng, Yingqi Zhao, Jiaheng Liu, Pengfei Gao, He Ye, Dayi Lin, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对编码智能体基准测试中持续执行洞察不足的问题,推出长周期基准 LoopsBench,含112个多领域任务,评估得最强配置解决25%任务,开源相关数据代码。

中文摘要 AI 辅助

随着编码智能体被部署用于持续的长周期软件开发,编码智能体基础设施正从 harness 工程转向 loop 工程。现有基准通常聚焦于局部任务或最终状态结果,对持续执行的洞察有限。我们推出 LOOPSBENCH,这是一个用于编码智能体评估的长周期 loop 工程基准。每个任务是一个依赖有向无环图(DAG),由可独立测试的开发单元构成,其先决条件边有来源证据支持。LOOPSBENCH 包含来自真实来源的 112 个任务,涵盖 8 种编程语言和 9 个领域。其感知流程的运行时会沿就绪前沿释放测试,并将已完成节点作为回归义务保留。我们评估了前沿编码智能体与广泛使用的 loop 实现的配对情况,最强配置为 Opus-4.7 搭配 Claude Code 及外部续行,可解决 25.00% 的任务。记录的计划仅恢复了部分来源恢复的先决条件 DAG,且在评估的 loop 配置中仍可见回归事件。我们在 microsoft/Loopsbench 开源了基准数据和代码,包括所有任务、5300 多个开发单元及可执行测试。

英文摘要

Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon software development. Existing benchmarks often center on localized tasks or end-state outcomes, offering limited insight into sustained execution. We introduce LOOPSBENCH, a long-horizon benchmark for loop engineering in coding agent evaluation. Each task is a dependency DAG over separately testable development units with source-evidenced prerequisite edges. LOOPSBENCH comprises 112 tasks from authentic sources spanning 8 programming languages and 9 domains. Its flow-aware runtime releases tests along the ready frontier and retains completed nodes as regression obligations. We evaluate frontier coding agents paired with widely used loop implementations. The strongest configuration, Opus-4.7 with Claude Code and outer continuation, resolves 25.00% of tasks. Recorded plans recover only part of the source-recovered prerequisite DAG, and regression events remain visible across the evaluated loop profiles. We open source the benchmark data and code, including all tasks, more than 5,300 development units, and executable tests, at microsoft/Loopsbench.

发表机构

  • Microsoft(微软)
  • Nanjing University(南京大学)
  • University College London(伦敦大学学院)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑