arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编码智能体驾驭(Harness)设计的实证研究

An Empirical Study of Harness Design for Coding Agents

Run-Ze Fan, Zihao Zhang, Simin Ma, Yebowen Hu, Shouju Wang, Kaiqiang Song, Fei Liu, Hamed Zamani, Xiaoyang Wang

arXiv 2609.20804首次发表:更新:

发表机构

UMass Amherst; Emory University; UNC Charlotte; Zoom Video Communications(马萨诸塞大学阿默斯特分校; 埃默里大学; 北卡罗来纳大学夏洛特分校; Zoom视频通信公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过轻量级编码驾驭系统,在四个模型和两个基准上实证分析了规划、动作空间与上下文管理三个组件的影响,发现上下文管理在预算紧张时防溢出最有效,规划随模型增强转为降本,动作空间影响成本与粒度,为驾驭设计提供模块化框架。

AI 中文摘要

编码驾驭(Coding harness)塑造了自主编码智能体将模型能力转化为长周期软件工程性能的方式,然而现有工作通常将驾驭系统作为整体进行评估,导致各独立组件的有效性尚不明确。为实现组件级比较,我们借助一个轻量级编码驾驭系统来研究该问题,其执行循环固定不变,而三个组件可进行变化:规划(planning)、动作空间(action space)和上下文管理(context management)。在SWE-Bench Verified和Terminal-Bench 2.1上评估的四个模型中,我们评估了176种匹配设置,涵盖五种上下文管理策略、四种上下文窗口预算,以及针对规划和动作空间的定向消融实验。我们发现:(1)随着上下文窗口预算收紧,上下文管理变得越来越有价值,其大部分收益来自防止上下文溢出(context-overflow)失败。(2)在基于LLM的摘要之前进行基于规则的省略(staging rule-based elision)在上下文管理策略中提供了最强的整体效率,而让省略内容可恢复则增加了模型很少使用的机制,且未带来准确率提升。(3)规划从较弱模型的准确率支架转变为较强模型的成本节约器,准确率变化很小。(4)对于bash熟练度较弱的模型,预定义工具可提升性能,而具备bash能力的模型可以仅通过bash接口有效运作,并实现显著更低的成本,尤其是在以命令行为中心的任务上。轨迹级分析解释了这些效应:上下文管理在不显著改变智能体行为的情况下延长了执行轨迹,规划改变了轨迹停止的位置,而动作空间则改变了代码编写的粒度。这些发现为模型感知和预算感知的驾驭设计提供了信息,并为评估未来驾驭组件提供了一个模块化框架。

英文摘要

Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management. Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we evaluate 176 matched settings spanning five context-management strategies, four context-window budgets, and targeted ablations of planning and action space. We find that: (1) Context management becomes increasingly valuable as the context-window budget tightens, with most of its benefit coming from preventing context-overflow failures. (2) Staging rule-based elision before LLM-based summarization provides the strongest overall efficiency among the context-management strategies, whereas making elided content recoverable adds machinery that models rarely use and yields no accuracy gain. (3) Planning shifts from an accuracy scaffold for weaker models to a cost saver for stronger models, with little change in accuracy. (4) Predefined tools improve performance for models with weaker bash proficiency, whereas bash-capable models can operate effectively with a bash-only interface and achieve substantially lower cost, especially on command-line-centric tasks. Trajectory-level analysis explains these effects: context management extends execution trajectories without substantially altering agent behavior, planning changes where trajectories stop, and the action space changes the granularity at which code is written. These findings inform model- and budget-aware harness design and provide a modular framework for evaluating future harness components.

Comments43 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑