arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26218cs.AIcs.SE

相同模型,不同工具链:编码智能体的结果差异

Same Model, Different Harness: Different Coding-Agent Results

Sydney Lewis

首次发表
浏览论文内容

中文总结 AI 辅助

本文探究模型与任务固定时,工具链改变对编码智能体结果的影响,发现特定工具链配置可提升模型在编码基准的表现,故编码智能体评估需将模型与工具链共同视为求解器。

中文摘要 AI 辅助

编码智能体将模型与工具链(harness)结合,工具链决定模型的可见内容、可使用的工具以及工作推进方式。本文在模型和任务固定的情况下,探究改变工具链是否会改变结果。我们在三个编码基准上对比了同一工具链的两种配置:对照组按时间顺序提供完整对话,实验组保留相同记录,但在上下文填满时机械缩短旧工具结果,并对重复或停滞的工作作出响应。在紧上下文设置下,实验组在全部三次压力对比中提高了平均每任务未通过分数(F2PF),并在SWE-bench Verified和SWE-bench Pro上增加了完整解决方案数量。紧窗口Verified对比使用169个任务、20480-token窗口和固定480秒尝试端点;在该组中,实验组将平均每任务F2PF从28%提升至49%,完整解决方案从43个提升至72个。无需针对模型进行特定微调,相同的冻结实验组还为三种不同设计的模型在同一组中同时提升了两个指标。在宽窗口Qwen3.6对比中,Verified和Pro上观察到的臂结果接近,而FeatureBench在实验组下保持更高的平均每任务F2PF;在宽窗口Verified组中,实验组每轮使用的提示token更少。由于改变工具链会改变未修改模型权重所能达成的效果,编码智能体评估应将模型和工具链共同视为被测试的求解器。

英文摘要

A coding agent combines a model with a harness, which decides what the model sees, which tools it can use, and how the work continues. We ask whether changing the harness changes the result when the model and task stay fixed. We compare two configurations of the same harness on three coding benchmarks. The control supplies the full conversation in time order, while the treatment keeps the same record but mechanically shortens older tool results as the context fills and responds to repeated or stalled work. Under tight context, the treatment raises mean per-task fail-to-pass fraction (F2PF) in all three pressure comparisons and increases complete solutions on SWE-bench Verified and SWE-bench Pro. The tight-window Verified comparison uses 169 tasks, a 20,480-token window, and a fixed 480-second attempt endpoint; on this cohort, treatment raises mean per-task F2PF from 28 percent to 49 percent and complete solutions from 43 to 72. Without model-specific retuning, the same frozen treatment also raises both endpoints on the same cohort for three additional models with different designs. In the wide-window Qwen3.6 comparisons, observed arm outcomes are close on Verified and Pro, while FeatureBench retains a higher mean per-task F2PF under treatment. On the wide-window Verified cohort, treatment also serves fewer prompt tokens per turn. Because changing the harness changed what unchanged model weights could accomplish, coding-agent evaluations should treat the model and harness together as the tested solver.

补充信息

↑