发表机构
PointFive(PointFive)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于 API 的编码代理上下文减少层,通过对 2848 次 Claude Code 运行分析,发现提示缓存流量主导成本,工具输出减少不能可靠预测计费成本降低,压缩可能损害任务完成,进而提出以成功调整计费成本为核心的分层证据标准。
AI 中文摘要
基于 API 的编码代理的上下文减少层,包括命令输出压缩器、检索排序器和有效载荷优化代理,通常通过它们去除的文本量来评估。我们提出相反的问题:在不降低任务成功率或延长其轨迹的情况下,何时减少检索到的上下文或工具输出会降低编码代理的实际计费成本?我们的主要证据是对 2908 次由提供商计费的 Claude Code 运行进行的预先指定、哈希冻结的配对活动,其中分析了 2848 次,涵盖 103 个任务、7 个存储库和 3 个模型。该活动在大约 5500 次计费执行的更广泛测量程序中,将基线与两代基于钩子的压缩和一个 API 边界代理进行了比较。出现了三个发现。首先,提示缓存流量主导成本构成。缓存创建和读取约占重建的四组件成本(约占实际账单的 80%)的 87%,有 8.7%的美元加权残余部分,保留的遥测数据无法归因。在 Haiku 4.5 上,此残余部分与思考工作量成比例。其次,工具输出减少并不能可靠地预测计费成本降低。一个去除了估计原始工具输出令牌 38%的分支,配对成本高出 6.8%(95%置信区间:+2.8%至+11.3%),而每任务减少与成本变化仅弱相关(皮尔逊 r = 0.15,置信区间跨越零)。第三,压缩可能会通过去除关键动作证据来损害任务完成。在一项关于源自 SWE-bench 的 Go 任务的小型单样本研究中,压缩通过破坏逐字编辑锚点将补丁应用从 27/40 减少到 15/40,并且压缩的基础分支在每个解决方案的观察成本较高时产生的解决方案更少。我们提出了一个分层证据标准,其核心是成功调整后的计费成本,而不是仅令牌减少。
英文摘要
Token-reduction tools for coding agents are often evaluated by the number of tokens they remove, but token count alone does not determine end-to-end inference cost. We evaluate three token-reduction approaches against an unmodified Claude Code baseline across controlled coding tasks, measuring provider-billed cost, task success, cache traffic, and agent behavior. The largest compression setup reduced delivered tool-output tokens by 38.4% but increased billed cost by 6.8%, while lighter compression produced only small and statistically uncertain savings. Across tasks, token reduction was weakly correlated with cost reduction (Pearson r = 0.15). Cost decomposition shows that prompt-cache creation and reads dominate the measured input-side cost, leaving only a limited fraction of total spend directly addressable by tool-output compression. We also find that compression can alter agent trajectories through additional retrieval, diagnosis, testing, and turns, offsetting local token savings. On a SWE-bench Go subset, aggressive compression also reduced successful patch application. These results show that token reduction is not a reliable proxy for cost reduction in tool-heavy coding agents. Effective optimization should therefore be evaluated at the level of cost per successful task, including cache behavior, trajectory changes, and correctness rather than token counts alone.