AI 中文总结
该研究通过对照实验发现,智能体脚手架对AI编码智能体的工具使用成本影响远大于接口,MCP并非必要,且智能体常忽略分配的接口,相关数据已开源。
AI 中文摘要
AI编码智能体的运行成本更多取决于驱动它的智能体脚手架,而非其访问工具的接口。我们计划测量通过模型上下文协议(MCP)使用工具与通过普通命令行界面(CLI)使用工具的成本,已发表的估算值相差超过一个数量级,且基于无法复现的从业者报告。我们针对一项固定软件任务——对私人在线Git仓库执行六项操作——在七种智能体脚手架和五种语言模型上运行,并通过检查仓库状态而非信任智能体的自我报告来验证完成情况。主导效应来自脚手架:七种脚手架中有两种完全不支持MCP,它们仅使用CLI完成所有运行,表明MCP对这类工作并非必要,且在不附加任何MCP服务器的情况下,仅比较CLI运行的成本,这两种脚手架比支持MCP的五种脚手架便宜5.0倍至28倍。对于本地运行的270亿参数小型模型,该效应最大:它在所有脚手架下均完成任务,且成本在不同脚手架间变化139倍。我们计划开展的比较被证明不稳定:13组严格配对的MCP与CLI成本比介于0.43倍至29倍之间,两侧均存在异常值。两种接口在失败成本上存在差异:MCP运行中花费的资金有12.9%未产生已完成的工作,而CLI运行中这一比例为2.2%;但失败频率无差异,原始运行及其重复运行中的失败情况同样常见。智能体经常忽略分配给它们的接口,因此未验证实际行为的比较测量的是未知混合结果。测试框架、任务、验证方式及完整数据集均以开源形式发布。
英文摘要
How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on the interface through which it reaches its tools. We set out to measure the cost of tool use over the Model Context Protocol (MCP) against tool use over an ordinary command-line interface (CLI), a difference on which published estimates disagree by more than an order of magnitude while resting on practitioner reports that cannot be reproduced. We ran one fixed software task -- six operations against a private online git repository -- across seven agent scaffoldings and five language models, and we verified completion by inspecting the repository state rather than trusting the agent's self-report. The dominant effect was the scaffolding. Two of the seven ship no MCP support at all; they completed every run using only the CLI, which shows that MCP is unnecessary for this class of work, and they were 5.0x to 28x cheaper than the five scaffoldings that do support MCP, comparing CLI runs alone with no MCP server attached anywhere. The effect was largest for a small 27-billion-parameter model running locally, whose cost varied 139x across scaffoldings while it completed the task under all of them. The comparison we set out to make proved unstable: thirteen strictly paired MCP-to-CLI ratios span 0.43x to 29x, with outliers on both sides. The two interfaces separate on the cost of failure, where 12.9 per cent of the money spent on MCP runs bought no completed work against 2.2 per cent on CLI runs, but not on its frequency: failures were equally common in both, in the original runs and in their repetitions alike. Agents frequently ignored the interface they were assigned, so comparisons that do not verify actual behaviour measure an unknown mixture. The harness, the task, the verification and the complete dataset are released as open source.
Comments29 pages, 4 figures, 8 tables. Companion methodology paper: arXiv:2606.11869. Dataset and measurement harness (open source, GPL-3.0): https://doi.org/10.5281/zenodo.21851992