ActKV:通过动作引导的KV缓存管理实现高效LLM智能体
ActKV: Efficient LLM Agents through Action-Guided KV Cache Management
浏览论文内容
中文总结 AI 辅助
ActKV提出面向智能体LLM推理的KV缓存压缩框架,通过动作导向驱逐、置信度驱动预算分配和页感知管理,在仅用25.98%内存时保留98.53%准确率,并提升3.97倍吞吐量。
中文摘要 AI 辅助
智能体LLM推理在迭代的观察-推理-动作循环中累积长KV缓存,造成显著的内存开销并限制服务吞吐量。现有压缩方法强调整体输出质量,忽视了动作在推动任务进展中的不对称重要性。我们的关键思想是建立一种压缩标准,根据KV条目对动作生成的贡献来评估其价值,并优先保证动作质量。然而,迭代执行、动态内存需求以及分散的动作关键条目对驱逐策略、预算分配和分页内存集成构成挑战。为此,我们提出ActKV,这是首个针对智能体LLM推理定制的KV缓存压缩框架。(i) 面向动作的KV缓存驱逐利用稳定的动作访问模式保留对未来动作至关重要的条目,支持压缩下的可靠任务进展。(ii) 置信度驱动的自适应预算分配利用LLM的内在置信度,使预算适应不断变化的动作关键内存需求。(iii) 页感知压缩管理将压缩标准化为三个原语,并配备定制内核,实现实际的吞吐量提升。在长轨迹任务上,ActKV仅使用FullKV峰值KV缓存内存的25.98%,平均保留了FullKV准确率的98.53%。它还实现了FullKV的3.97倍和3.58倍的令牌和任务吞吐量,提供了最先进的性能。
英文摘要
Agentic LLM inference accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput. Existing compression methods emphasize overall output quality, overlooking the asymmetric importance of actions in driving task progress. Our key idea is to establish a compression criterion that values KV entries by their contribution to action generation and prioritizes action quality. However, iterative execution, dynamic memory demands, and scattered action-critical entries pose challenges to eviction policies, budget allocation, and paged memory integration. To this end, we propose ActKV, the first KV cache compression framework tailored for agentic LLM inference. (i) Action-oriented KV cache eviction exploits stable action access patterns to retain entries critical to future actions, supporting reliable task progress under compression. (ii) Confidence-driven adaptive budget allocation uses LLM's intrinsic confidence to adapt the budget to evolving action-critical memory demands. (iii) Page-aware compression management standardizes compression into three primitives with customized kernels, realizing practical throughput gains. On long-trace tasks, ActKV retains an average of 98.53% of FullKV's accuracy with only 25.98% of its peak KV cache memory. It also achieves 3.97 times and 3.58 times FullKV's token and task throughput, delivering state-of-the-art performance.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- Suzhou Institute for Advanced Research, University of Science and Technology of China(中国科学技术大学苏州高等研究院)
机构由 AI 辅助整理,请以论文原文为准。