发表机构
Shandong University(山东大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过固定模型、更换 harness 的对照实验,量化了编码智能体 harness 对性能的影响,发现其效果常被噪声淹没,但显著影响成本,且成本差异可达 3 倍。
AI 中文摘要
编码智能体是一个由“harness”包裹的语言模型:系统提示、工具集和上下文管理将聊天模型转变为能够在代码仓库中工作的系统。生产环境的 harness 每天发布版本,供应商宣传 harness 变更带来的通过率提升,排行榜也自由地混合使用不同的 harness。但很少被衡量的是,当模型固定不变时,harness 本身对分数的影响有多大。我们在 SWE-bench Verified 上,通过三个生产级 harness(Claude Code、mini-SWE-agent 和 OpenCode)运行了五个模型,并重复运行相同的配置,以校准当除运行本身外其他条件不变时分数变化的幅度。在 447 个任务上,对于我们所运行的两个模型(Claude Code 和 mini-SWE-agent),最重的和最轻的 harness 之间的分数差异在 5 分以内。在一个包含 45 个任务的困难子集上,使用五个模型,更换 harness 所翻转的任务数量与重复运行同一 harness 所翻转的任务数量相当,两者均为 13%,而且一个 harness 在某次运行中获胜的任务,在下次运行中并不一定仍然获胜。唯一能超越噪声的 harness 效应是损失而非收益:OpenCode 在大规模任务池上落后多达 9 分,而在其中一个模型上,这一差距的一半来自于其输出上限截断的运行。harness 真正决定的是成本。在相同模型、相同任务和相同价格表下,不同 harness 的单任务成本差异高达 3 倍。这一差距在第一次调用时就已确定,由每个 harness 每一步发送的系统提示和工具模式的序言决定,并随步数缩放;每步的增长和每次调用的工具输出差异要小得多。提供商对缓存输入的定价会缩放账单,但不会改变其顺序。重复运行的数据也提供了 harness 比较所需的分辨率:在我们观察到的 discordance 下,45 个任务只能以一半的概率捕捉到 13 分的差距,且无法以 80% 的统计功效检测到任何差距;447 个任务可以分辨 5 分的差距,但仍比许多 harness 变更声称的收益要粗糙。
英文摘要
A coding agent is a language model wrapped in a harness: the system prompt, the tool set, and the context management that turn a chat model into something that can work inside a repository. Production harnesses ship releases daily, vendors advertise pass-rate gains from harness changes, and leaderboards mix harnesses freely. What is rarely measured is how much the harness itself moves the score when the model is held fixed. We run five models through three production harnesses, Claude Code, mini-SWE-agent, and OpenCode, on SWE-bench Verified, and rerun the same configurations to calibrate how much a score moves when nothing changes but the run. On 447 tasks and the two models we ran there, Claude Code and mini-SWE-agent, the heaviest and the lightest harness, are equivalent within five points. On a 45-task hard subset and five models, swapping the harness flips as many tasks as rerunning the same harness, 13% in both cases, and the tasks a harness wins in one run are not the tasks it wins in the next. The one harness effect that clears the noise is a loss, not a gain: OpenCode trails by up to 9 points on the large pool, and on one model half of that gap sits in runs its output cap cut short. What the harness does decide is the bill. With the same model, the same tasks, and one price list, cost per task differs by up to 3x across harnesses. The gap is set at the first call, by the preamble of system prompt and tool schemas each harness sends with every step, and scaled by the number of steps; per-step growth and per-call tool output differ far less. The provider's price for cached input scales the bill and does not reorder it. The rerun data also give the resolution a harness comparison needs: at the discordance we observe, 45 tasks catch a 13-point gap only half the time and no gap with 80% power, and 447 tasks resolve 5 points, still coarser than the gains many harness changes claim.
Comments22 pages, 7 figures. Code and data: https://github.com/YangzeLiu/what-does-a-harness-buy