ReCache:面向工具增强型大语言模型智能体的高效键值缓存复用与压缩
ReCache: Efficient KV Cache Reuse and Compression for Tool-Augmented LLM Agents
- Shanghai Jiao Tong University(上海交通大学)
- Eastern Institute of Technology(东方理工学院)
- Xi’an Jiaotong University(西安交通大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
ReCache是面向工具增强型大语言模型智能体的框架,通过资源级注意力与剪枝实现KV缓存复用压缩,在性能损失极小的情况下大幅降低推理计算与内存开销,已在多工具数据集基准上验证有效。
AI中文摘要:
智能体语言模型会反复对不同请求中以不同组合和顺序出现的工具与技能模式进行编码,这使得标准前缀缓存无法复用它们的键值(KV)状态。本文提出ReCache,一种用于独立缓存资源表示同时降低推理时计算与内存开销的框架。资源级注意力消除了资源间的交互,并为资源分配局部位置,生成具有组合不变性的KV块。ReCache随后将资源可见性限制为经贡献选择的层-KV头组路由,并通过结构和语义剪枝仅保留调用关键字段。我们在由7个公开工具与技能使用数据集组成的基准上评估ReCache,包括资源不相交测试。资源级注意力的调用性能与密集模型相当(Inv-F1分别为82.3%与82.4%),同时实现了3.655倍的首token生成速度提升。完整框架将分配的KV张量内存降低92.43%,并将注意力计算速度提升1.423倍。这些结果表明,将可复用模式编码与选择性资源访问分离,可在有效性损失有限的情况下大幅降低智能体推理成本。代码可在该https网址获取。
英文摘要:
Agentic language models repeatedly encode tool and skill schemas that recur across requests in different combinations and orders, preventing standard prefix caching from reusing their key--value (KV) states. We introduce \textbf{ReCache}, a framework for independently caching resource representations while reducing their inference-time computational and memory overhead. Resource-wise attention removes cross-resource interactions and assigns resource-local positions, producing composition-invariant KV blocks. ReCache then restricts resource visibility to contribution-selected layer--KV-head-group routes and retains only invocation-critical fields through structural and semantic pruning. We evaluate ReCache on a benchmark assembled from seven public tool- and skill-use datasets, including resource-disjoint tests. Resource-wise attention matches dense invocation performance (82.3\% versus 82.4\% Inv-F1) while providing a 3.655$\times$ time-to-first-token speedup. The complete framework reduces allocated KV-tensor memory by 92.43\% and accelerates attention by 1.423$\times$. These results show that separating reusable schema encoding from selective resource access substantially reduces agentic inference costs with limited effectiveness loss. The code is available at https://github.com/EIT-NLP/ReCache.