arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22158cs.LG

StepKV:面向LLM智能体的步骤感知KV缓存压缩

StepKV: Step-Aware KV Cache Compression for LLM Agents

  • The Chinese University of Hong Kong(香港中文大学)
  • Huawei Technologies Co., Ltd(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Boyu Feng, Jiahong Liu, Yifan Li, Wenhao Yu, Zexuan Qiu, Yuliang Sun, Ming Shen, Xiang Li, Quanyu Dai, Irwin King

AI总结:

StepKV提出以推理步骤为保留单元的KV缓存压缩方法,结合步骤效用与令牌显著性,在多跳问答和长时程网络推理中于低预算下保持准确率,优于令牌级基线。

AI中文摘要:

键值(KV)缓存对于高效的自回归大语言模型(LLM)推理至关重要,但缓存随上下文长度线性增长,增加了存储和解码成本。KV缓存压缩通过仅保留缓存令牌的子集来缓解这一成本。这一挑战对于多步骤LLM智能体尤为重要,因为查询会扩展为推理轨迹、工具交互和检索到的观察结果。现有的剪枝方法通常将缓存视为扁平令牌流,并按近期性或注意力显著性对令牌进行排序。这造成了压缩单元与推理单元之间的不匹配:令牌级剪枝移除单个条目,而多步骤智能体中的有用信息通常组织为推理步骤,这些步骤的重要性不均匀且延迟显现。因此,早期的观察或中间决策可能很少受到近期关注,但对于后续的证据综合仍然至关重要。我们将这种失败模式称为“推理连续性”。这些观察促使KV缓存压缩同时考虑令牌级和推理步骤级信息。StepKV通过将推理步骤视为第一类保留单元来实现这一目标。它将缓存条目与其生成步骤关联,从轨迹派生信号估计步骤效用,并将该效用与令牌级显著性相结合。所得分数对可剪枝令牌进行全局排序,StepKV在目标预算下保留得分最高的条目。因此,StepKV为智能体KV缓存压缩提供了以步骤为中心的视角。在多跳问答和长时程网络推理任务中,StepKV在低KV预算下保持了准确性,而令牌级基线则急剧下降,为多步骤智能体推理提供了更稳健的效率-准确性权衡。

英文摘要:

Key-value (KV) caching is essential for efficient autoregressive large language model (LLM) inference, but the cache grows linearly with context length, increasing storage and decoding costs. KV cache compression mitigates this cost by retaining only a subset of cached tokens. This challenge is particularly important for multi-step LLM agents, where a query expands into trajectories of reasoning, tool interactions, and retrieved observations. Existing pruning methods typically treat the cache as a flat token stream and rank tokens by recency or attention saliency. This creates a mismatch between the unit of compression and the unit of reasoning: token-level pruning removes individual entries, whereas useful information in multi-step agents is often organized into reasoning steps with uneven and delayed importance. Consequently, an early observation or intermediate decision may receive little recent attention yet remain essential for later evidence synthesis. We term this failure mode Reasoning Continuity Disruption.These observations motivate KV cache compression that jointly considers token- and reasoning-step-level information. StepKV addresses this goal by treating reasoning steps as first-class retention units. It associates cache entries with their generating steps, estimates step utility from trajectory-derived signals, and combines this utility with token-level saliency. The resulting scores globally rank prunable tokens, from which StepKV retains the top-scoring entries under a target budget. StepKV thus provides a step-centric perspective for agent KV cache compression. Across multi-hop QA and long-horizon web reasoning tasks, StepKV sustains accuracy under low KV budgets where token-level baselines degrade sharply, offering a more robust efficiency-accuracy trade-off for multi-step agent inference.

↑