arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编码智能体中的支架效应:在编码智能体评估中利用选择作为隐藏变量

The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation

Naman Vats, Oleg Golev

arXiv 2607.22585首次发表:更新:

发表机构

Alibaba; Agentic AI Foundation; OpenCode Contributors; Harbor Framework Team(阿里巴巴; Agentic人工智能基金会; OpenCode贡献者团队; Harbor框架团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究编码智能体评估中工具选择对结果的影响,通过在三个开源工具上评估两个模型,发现工具选择会使解决任务的令牌数有40倍差异,失败指纹有工具级偏差,建议按通过率选工具-模型对并报告相关指标,还发布了相关数据和脚本。

AI 中文摘要

编码智能体的公共排行榜通常按模型名称和通过率对系统进行排名,而周围的工具(发出工具、管理上下文并决定何时停止的支架)往往规定不明确。当工具固定时,模型间比较有效;当工具变化时,性能和效率会混淆模型和支架效应。我们在Terminal-Bench Pro的一个分层50任务子集上,通过三个开源工具(Goose、OpenCode、OpenHands-SDK)评估了Qwen 3.6 Plus和MiniMax M2.5。工具选择会导致每个解决任务的令牌数相差40倍,而模型内配对的通过率差异保持在0-8个百分点。失败指纹在不同模型间重复出现,表明工具级偏差在很大程度上与模型无关。对于以人类为中心的编码智能体评估,仅模型名称是不完整的比较单元,工具-模型对决定实际成本、延迟和监督负担。因此建议在令牌/延迟预算下按通过率选择工具-模型对,并在任何模型比较时报告令牌使用情况、延迟和完整的工具规范。我们还发布了匿名配置、原始试验日志、聚合快照和分析脚本。

英文摘要

Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools, manages context, and decides when to stop) is often under-specified. Model-to-model comparison is valid when the harness is fixed; when it varies, performance and efficiency conflate model and scaffold effects. We evaluate Qwen 3.6 Plus and MiniMax M2.5 across three open-source harnesses (Goose, OpenCode, OpenHands-SDK) on a stratified 50-task subset of Terminal-Bench Pro. Harness choice induces up to a 40x difference in tokens per solved task, while paired within-model pass-rate differences remain 0-8 percentage points (95% paired-task bootstrap CIs include zero except for the largest gap). Failure fingerprints replicate across models (REASON for Goose, VERIFY/MAX_TURNS for OpenHands-SDK, idle-loop/TIME for OpenCode), indicating harness-level biases that are largely model-independent. For human-centered coding-agent evaluation, model name alone is an incomplete comparison unit: harness-model pairs determine real-world cost, latency, and oversight burden; no-action turns are a per-task wait tax, not just a token tax. We therefore recommend selecting harness-model pairs by pass rate under token/latency budgets, and reporting token usage, latency, and full harness specifications alongside any model comparison. We release anonymized configs, raw trial logs, aggregated snapshots, and analysis scripts.

Comments7 pages, 2 figures, 6 tables. Preliminary work; under review at the 5th DL4C Workshop @ ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑