多轮编码智能体中上下文压缩网关的经验成本归因
An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents
- Paritok
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
通过分解多轮编码智能体中压缩网关的令牌账单,发现工具模式过滤线性节省成本,内容压缩二次方节省但受上下文限制,召回成本有界,单次压缩基准不能证明多轮成本节省。
AI中文摘要:
上下文压缩被广泛提出作为削减LLM编码智能体令牌账单的一种方式,公开基准报告称激进的压缩能保持任务解决质量。这两个事实并不暗示通常假设的第三个事实:压缩文件读取在多轮真实智能体中能节省成本。我们在编码智能体(Claude Code、Codex)与前沿LLM(Claude Sonnet、GPT-5)之间对一个生产压缩网关(Paritok)进行了插桩,并将真实会话的令牌账单分解为三个独立的杠杆:工具模式过滤、文件读取和工具输出的内容压缩,以及历史摘要。在受控A/B运行中单独测量时,这三者的节省速率根本不同。工具模式过滤每轮移除一个固定块,在典型轮次中约为21K-57K个令牌;它随轮次N线性增长,是唯一明确且可复现的正向杠杆。内容压缩每轮仅节省约2%的缓存定价前缀,但压缩的读取会累积在历史中并在后续每轮重新发送,因此其累积节省呈二次方增长,约3350*N^2个令牌(实测),在大约6轮内超过固定工具过滤的节省,直到上下文窗口将其封顶。非破坏性网关允许智能体按需拉回原始字节;每次召回恰好重新发送刚被压缩掉的一个片段,因此其成本固定且有界,而非乘法式爆炸,重度召回会逐片段花掉累积的节省。最后,一个强大的单次压缩基准——在25.7%压缩率下保留86.5%的SWE-bench质量,由该网关部署的模型(Paritok-4B,另行报告)实现——与多轮智能体成本正交,不得引用为节省成本的论据。我们将结果提炼为一份可操作的指南,说明令牌节省工作在哪里能产生回报。
英文摘要:
Context compression is widely proposed as a way to cut the token bill of LLM coding agents, and public benchmarks report that aggressive compression preserves task-solving quality. These two facts do not imply the third one commonly assumed: that compressing file reads saves money in a real multi-turn agent. We instrument a production compression gateway (Paritok) between coding agents (Claude Code, Codex) and frontier LLMs (Claude Sonnet, GPT-5), and decompose the token bill of real sessions into three independent levers: tool-schema filtering, content compression of file reads and tool output, and history summarization. Measured in isolation under controlled A/B runs, the three save at fundamentally different rates. Tool-schema filtering removes a fixed block every turn, roughly 21K-57K tokens on a typical turn; it is linear in the turn count N and the only unambiguously and reproducibly positive lever. Content compression saves only about 2% of the cache-priced prefix per turn, but compressed reads accumulate in history and are re-sent on every later turn, so its cumulative saving grows quadratically, about 3350*N^2 tokens (measured), overtaking the fixed tool-filter saving within roughly 6 turns until the context window caps it. A non-destructive gateway lets the agent pull original bytes back on demand; each recall re-sends exactly the one segment just compressed away, so its cost is fixed and bounded rather than a multiplicative blowup, and heavy recall spends the accumulated saving back one segment at a time. Finally, a strong single-shot compression benchmark - 86.5% of SWE-bench quality retained at a 25.7% compression rate, achieved by the model this gateway deploys (Paritok-4B, reported separately) - is orthogonal to multi-turn agent cost and must not be cited as a cost-saving argument. We distill the results into an actionable recipe for where token-saving effort pays off.