arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11987cs.AIcs.CLcs.SE

驾驭还是模型?在污染受控的私有测试集上隔离智能体编码中的驾驭效应

Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite

  • evolutionID GmbH(evolutionID 有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Mohsen Arjmandi

AI总结:

通过污染受控的私有测试集,对比原生与中性驾驭在智能体编码中的表现,发现无平均优势,但任务类型影响显著,且成本差异受数据缺失影响。

AI中文摘要:

智能体编码系统将语言模型与驾驭(harness)耦合在一起:驾驭包括工具、提示词和控制流,它们将聊天模型转变为自主软件工程师。供应商会为其自有模型调校驾驭,而从业者假定供应商原生的配对能解决更多任务。我们在一个包含256个仓库任务和截止日期后竞赛任务的私有、污染受控测试集上,通过配对的同模型对比来衡量这一假设。相同的80个任务在claude-agent-sdk和deepagents上运行于claude-opus-4-8,在openai-codex SDK和deepagents上运行于gpt-5.5,并以gemini-3.5-flash和deepseek-v3.2作为辅助单元。800个计划运行中的792个由隔离的预言机(oracle)评分。两个对比均未显示出任一驾驭的平均优势:Opus 4.8为-1.25个百分点(48.8%对50.0%,任务自助法95%置信区间[-10.0, +7.5]),GPT-5.5为+1.25个百分点(55.6%对54.4%,置信区间[-4.4, +6.9])。Opus的平均值结合了相反的层次:在61个仓库任务上,原生驾驭落后9.0个百分点,而在19个竞赛任务上领先23.7个百分点(标签置换p=0.003)。该划分是在查看数据后选择的,需要设计好的重复实验。正确性和完成度也出现分离:在墙钟时间上限取消的81次运行中,有22次已产生通过补丁。按冻结标价从原始每轮使用量重新计价,在Opus 4.8上,中性驾驭每个已解决任务的成本为1.3至1.6倍,在GPT-5.5上为1.2倍。这些是观察到的使用量估计。在Anthropic账户上,58次运行未留下使用记录,将该支出分配给任一单元会使Opus比率在0.7至2.3之间移动,因此计费顺序尚未确定。本修订版更正了2026年8月的手稿,其成本数据基于我们自身遥测中的使用语义缺陷(第5.1节)。我们发布编排器、评分预言机、重分析代码和派生聚合数据。任务保持私有。

英文摘要:

An agentic coding system couples a language model to a harness: the tools, prompts and control flow that turn a chat model into an autonomous software engineer. Vendors ship harnesses tuned to their own models, and practitioners assume the vendor-native pairing solves more tasks. We measure that assumption with paired same-model contrasts on a private, contamination-controlled suite of 256 repository and post-cutoff contest tasks. The same 80 tasks ran under claude-agent-sdk and under deepagents on claude-opus-4-8, and under the openai-codex SDK and deepagents on gpt-5.5, with gemini-3.5-flash and deepseek-v3.2 as side cells. 792 of 800 planned runs were graded by an isolated oracle. Neither contrast resolves an average advantage for either harness: -1.25 pp for Opus 4.8 (48.8% vs 50.0%, task-bootstrap 95% CI [-10.0, +7.5]) and +1.25 pp for GPT-5.5 (55.6% vs 54.4%, CI [-4.4, +6.9]). The Opus average combines opposite strata: the native harness trails by 9.0 pp on the 61 repository tasks and leads by 23.7 pp on the 19 contest tasks (label-permutation p = 0.003). The partition was chosen after seeing the data and needs a designed replication. Correctness and completion also separate: 22 of 81 runs cancelled at the wall-clock ceiling had produced a passing patch. Re-priced from raw per-turn usage at frozen list prices, the neutral harness cost 1.3 to 1.6 times as much per solved task on Opus 4.8 and 1.2 times on GPT-5.5. These are observed-usage estimates. On the Anthropic account 58 runs left no usage record, and allocating that spend to either cell would move the Opus ratio between 0.7 and 2.3, so the billed ordering is unresolved. This revision corrects an August 2026 manuscript whose cost figures rested on a usage-semantics defect in our own telemetry (Section 5.1). We release the orchestrator, grading oracle, reanalysis code and derived aggregates. The tasks stay private.

补充信息

↑