arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26779cs.AIcs.LGcs.SE

CliffCompaction:面向长时程编码智能体的成本高效压缩

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

发表机构卡内基梅隆大学 · 博世人工智能中心
查看机构详情
  • Carnegie Mellon University(卡内基梅隆大学)
  • Bosch Center for AI(博世人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

Trang Nguyen, Eulrang Cho, Bingqing Chen, Tim Dettmers

首次发表
浏览论文内容

中文总结 AI 辅助

提出CliffCompaction自动压缩技术,通过仅截断不重写保持信息忠实,降低长时程编码智能体成本达50%,提升测试时扩展效率,并在KernelBench上超越专门方法。

中文摘要 AI 辅助

智能体通常需要处理需要数百万token上下文的复杂问题,由于上下文窗口有限,跨会话时必须进行压缩。我们开发了CliffCompaction,一种自动压缩技术,在有界上下文下将成本降低高达50%,同时在Terminal-Bench上保持或提升性能,并为测试时扩展实现新的效率水平,在KernelBench上取得最先进结果。CliffCompaction每次rollout的节省使测试时扩展的性能-成本权衡更加高效,以低于两次全上下文运行的成本,在Terminal-Bench上增加了超过10个百分点。在并行测试时扩展下,CliffCompaction使Kimi K2.6匹配Opus 4.7,并以更低成本超越Opus 4.6和GPT-5.3 Codex。CliffCompaction有效性的关键在于,它通过仅截断或丢弃内容来保持压缩信息的忠实性,从不改写或重写。我们从不压缩压缩结果——每次操作仅针对原始内容,先前压缩的输出被丢弃,防止上下文漂移累积。这些特性支持在超过百万token的会话中持续学习:在KernelBench上,CliffCompaction在200步后达到CUDA内核加速2.23倍,在400步后达到3.58倍,尽管是一种通用压缩技术,却超越了专门的搜索算法和训练过的智能体。我们开源了一个与脚手架无关的API代理实现,可用于Claude Code、Codex及其他框架。

英文摘要

Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on Terminal-Bench and achieving new levels of efficiency for test-time scaling and state-of-the-art results on KernelBench. The per-rollout savings of CliffCompaction make the performance--cost trade-off of test-time scaling more efficient, adding over 10 percentage points on Terminal-Bench for less than the cost of two full-context runs. Under parallel test-time scaling, CliffCompaction lets Kimi K2.6 match Opus 4.7, and exceed Opus 4.6 and GPT-5.3 Codex at lower cost. The key to CliffCompaction's effectiveness is that it keeps compacted information faithful by only truncating or dropping content, never rephrasing or rewriting it. We never compact a compaction---each pass operates only on original content, and prior compacted output is discarded, preventing context drift from accumulating. These properties sustain continual learning over sessions exceeding a million tokens: on KernelBench, CliffCompaction reaches CUDA kernel speedups of $2.23\times$ after 200 steps and $3.58\times$ after 400 steps, surpassing specialized search algorithms and trained agents despite being a general-purpose compaction technique. We open-source a scaffold-agnostic API-proxy implementation of CliffCompaction usable with Claude Code, Codex and other harnesses.

↑