发表机构
Independent Researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型智能体KV缓存内存瓶颈问题,提出无需训练的区域感知逐出策略MemDecay,为令牌分配特定区域优先级和衰减率,经实验验证该策略能有效管理缓存,确立语义提示结构对KV缓存管理的作用并明确其与关注重要性的结合方式。
AI 中文摘要
大语言模型智能体积累了异构上下文,其键值(KV)缓存可能成为主要内存瓶颈。现有逐出策略对每个令牌应用相同规则,忽略了编排器可用的语义结构。我们引入MemDecay,一种无需训练的区域感知KV缓存逐出策略。它为令牌分配特定区域的基本优先级和衰减率,在令牌获得关注时刷新保留分数,在固定缓存预算下逐出得分最低的页面,同时允许固定关键区域。我们还提供了从测量的关注寿命校准衰减率的过程。使用Qwen2.5 - 1.5B和3B在约450和1700令牌上下文下评估MemDecay。结果表明,跨区域关注寿命相差一个数量级,固定关键区域能保持系统区域事实的全缓存准确性,区域感知保留在上下文增长时依然有效,而基于近期性的保留则失效。消融实验确定关注分数归一化是当前公式的主要限制。这些结果确立了语义提示结构作为KV缓存管理的强大信号,同时阐明了应如何将其与基于关注的重要性相结合。
英文摘要
Large language model (LLM) agents accumulate heterogeneous context, including system instructions, plans, user turns, retrieved documents, tool outputs, and intermediate reasoning, whose key-value (KV) cache can become a major memory bottleneck. Existing eviction policies generally apply the same attention- or recency-based rule to every token, ignoring semantic structure already available to the agent orchestrator. We introduce MemDecay, a training-free, region-aware KV-cache eviction policy. MemDecay assigns tokens region-specific base priorities and decay rates, refreshes retention scores when tokens receive attention, and evicts the lowest-scoring pages under a fixed cache budget while allowing critical regions to be pinned. We also provide a procedure for calibrating decay rates from measured attention lifetimes. We evaluate MemDecay at approximately 450 and 1,700 token contexts using Qwen2.5-1.5B and 3B. Across all settings, attention lifetimes differ by an order of magnitude across regions: system-token half-lives range from 148 to 189 decoding steps, compared with 14 to 16 for scratchpad tokens. Pinning preserves system-region facts at full-cache accuracy in every setting, while no baseline preserves more than 13 of 24. Region-aware retention remains effective as context grows, whereas recency-based retention collapses. Accumulated-attention retention performs better on unpinned content, however, and ablations identify attention-score normalization as the main limitation of the current formulation. These results establish semantic prompt structure as a robust signal for KV-cache management while clarifying how it should be combined with attention-based importance.