arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReFold:面向长时程智能体的免训练可逆回合间上下文折叠

ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

Yupeng Su, Jiayi Tian, Zheng Zhang, Souvik Kundu

arXiv 2610.07863首次发表:更新:

发表机构

University of California, Santa Barbara; Intel(加州大学圣塔芭芭拉分校; 英特尔)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ReFold是一种免训练的渲染层方法,通过可逆地折叠回合间冗余内容来压缩长时程智能体的上下文,在保持任务成功率的同时显著降低token消耗、内存和推理成本。

AI 中文摘要

长时程LLM智能体基于仅追加的交互历史进行决策,该历史在每一步都会重新发送给模型,因此上下文及其成本随步骤增加而增长,直至会话超出上下文窗口。现有方法通过上下文需求预测来管理上下文,依赖额外的模型调用、启发式规则或训练策略。然而,这些预测性方法引入了运行时开销,使前缀缓存失效,并永久丢弃内容且无法保证恢复。为克服这些局限,我们提出ReFold:一种免训练的渲染层,它保留底层交互历史,同时仅压缩模型渲染的上下文。它无需辅助预测器即可消除两类回合间冗余:先前回合已显示的内容(替换为占位符),以及智能体自身报告已完成并折叠为一行注释的回合。两种操作均采用分块渲染,每几步而非每一步重写缓存前缀。每次移除都是严格可逆的,错误的移除仅需从历史中恢复一次,而非永久丢失内容。由于ReFold在渲染层操作,因此可即插即用地应用于标准ReAct式框架。在五个长时程基准和两个前沿LLM上的评估表明,ReFold在不降低任务成功率的情况下,将token消耗减少高达2.5倍,并将每个会话的KV缓存内存减半。在受限上下文预算下,它避免了高达92%的强制压缩。在并发服务负载下,它将请求排队延迟减少高达100%,推理加速高达1.7倍,同时推理成本降低高达3.4倍。

英文摘要

Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window. Existing methods manage the context through context requirement prediction, relying on additional model calls, heuristic rules, or trained policies. However, these predictive approaches introduce runtime overhead, invalidate prefix caches, and permanently discard content with no guarantee of recovery. To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context. It removes two kinds of inter-turn redundancy without an auxiliary predictor: content an earlier turn already displayed, replaced by a stub, and turns the agent itself reports finished, folded into a one-line note. Both operators use chunked rendering, rewriting the cached prefix once every few steps rather than at every step. Every removal is strictly reversible, a wrong removal costs one restore from the history rather than permanent content loss. Because it operates at the rendering layer, ReFold is plug-and-play across standard ReAct-style harnesses. Evaluations across five long-horizon benchmarks and two frontier LLMs demonstrate that ReFold reduces token consumption by up to 2.5x and halves the KV-cache memory per session without degrading task success rates. Under capped context budgets, it avoids up to 92% of forced compactions. Under concurrent serving workloads, it reduces request queuing delays by up to 100%, accelerating inference by up to 1.7x, while cutting inference costs by up to 3.4x.

Comments27 pages, 6 figures, 14 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑