arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10569cs.SEcs.AIcs.LG

限制编码代理执行代码何时有用?一种机制×代理设计消融研究

When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation

Hong Yang, Qi Yu, Travis Desell

首次发表
浏览论文内容

中文总结 AI 辅助

研究现代编码代理不同工具表面有效性,通过在特定任务上对两种代理进行三臂消融实验,发现最便宜工具表面由任务机制和代理设计共同决定,主要成本信号是缓存调整成本而非通过率。

中文摘要 AI 辅助

现代编码代理具有多种工具表面,如IDE原语、bash和模型上下文协议(MCP)代码执行,对此领域有三种相互矛盾的说法。我们进行了缺失的交叉比较,在合成计算任务和SWE-bench Mini修改任务上进行完整性良好的三臂消融实验(基线/bash_only/code_only),固定模型、框架和提示,使用两种代理(Claude Code、OpenAI Codex CLI),跨越机制和代理设计轴。结果显示,在四个(机制,代理)单元中,将代理限制为单个执行代码MCP工具在三个单元中比其最便宜的丰富工具对手更便宜或在统计上相当。唯一例外是SWE-bench/Claude,code_only在成本上有方向性增加但不显著。这表明最便宜的工具表面由任务机制和代理设计共同决定,且主要成本信号存在于缓存调整成本中而非通过率。

英文摘要

Modern coding agents expose multiple tool surfaces -- IDE primitives, bash, and Model Context Protocol (MCP) code-execution -- and the field has shipped three contradictory claims about which one matters. We run the missing crossed comparison: an integrity-clean three-arm ablation (baseline / bash_only / code_only) on synthetic computation tasks and SWE-bench Mini modification tasks, holding model, harness, and prompts fixed, with two agents (Claude Code, OpenAI Codex CLI) so the comparison spans both regime and agent-design axes. Across the four resulting (regime, agent) cells, restricting the agent to a single execute_code MCP tool is cheaper than -- or statistically tied with -- its cheapest tool-rich rival in three cells (significantly on Artifact/Claude and SWE-bench/Codex; directionally on Artifact/Codex), with pass rates statistically tied within each cell. The lone exception is SWE-bench/Claude, where code_only is directionally costlier (+14.4%, not significant); a conditional-cost analysis localizes that gap to failure-cost on doomed-run trajectories, not a per-edit tax on successful runs. Two implications: the cheapest tool surface is jointly determined by task regime and agent design rather than by either axis alone, and the headline cost signal lives in cache-adjusted cost -- not pass rate, which is invariant across surfaces at the model sizes we evaluate. The benchmark harness, task suite, and analysis code are available at https://github.com/hyang0129/onlycodes.

发表机构

  • Rochester Institute of Technology(罗切斯特技术学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑