arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30310cs.LGcs.AI

Tail-Replay:在混合大语言模型的前缀缓存中规避线性注意力的线性诅咒

Tail-Replay: Escaping the Curse of Linear Attention in Prefix Caching for Hybrid LLMs

  • Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信人工智能研究院(TeleAI))
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Yirui Liu, Ruoling Qi, Xuaner Wu, Penghang Liu, Jian Chen

AI总结:

Tail-Replay是一种混合大语言模型的前缀缓存机制,通过重放匹配前缀的短近期后缀重构线性注意力状态,实现无约束令牌级前缀复用,在低重放预算下保持高预填充质量并显著提升长前缀服务效率。

AI中文摘要:

混合大语言模型通过将全注意力层与线性注意力层交错排列,以降低长上下文推理的成本。这种结构使前缀缓存变得复杂:全注意力的键值缓存是可按令牌寻址的,而线性注意力层维护的循环状态无法回滚到任意前缀边界。现有的混合前缀缓存方法通过存储循环状态检查点来解决这种不匹配问题,导致令牌级匹配仅在与存储的检查点对齐的位置直接可用,将前缀复用限制在一组离散的边界内。本文提出Tail-Replay,一种能在混合大语言模型中实现无约束令牌级前缀复用的前缀缓存机制。核心见解是,诸如Gated DeltaNet之类的线性注意力机制可视为对输入前缀的结构化有损压缩:门控循环更新会逐步衰减早期输入的贡献。因此,匹配前缀的循环状态可通过仅重放该前缀的短近期后缀来很好地近似。Tail-Replay利用这一特性,仅缓存精确的全注意力键值缓存,而省略循环状态检查点。在缓存命中时,它通过重放匹配前缀的短近期后缀来重构线性注意力状态,从而使复用边界由共享令牌而非循环状态检查点决定。我们在三个基于Gated DeltaNet的混合模型上,使用LongBench和RULER基准评估Tail-Replay。仅需5%-10%的重放预算,它在LongBench和RULER上即可保留92.8%-99.9%的全预填充质量。为评估服务效率,我们在8K、16K和32K等多个匹配前缀长度下评估首令牌生成时间的加速比,加速比随前缀长度增加而增长,在32K时达到全预填充的9.1-14.3倍。

英文摘要:

Hybrid large language models interleave full-attention layers with linear-attention layers to reduce the cost of long-context inference. This structure complicates prefix caching: full-attention key-value caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbitrary prefix boundaries. Existing hybrid prefix caching methods address this mismatch by storing recurrent-state checkpoints. As a result, token-level matches are directly usable only at positions aligned with stored checkpoints, constraining prefix reuse to a discrete set of boundaries. We present Tail-Replay, a prefix caching mechanism that enables unconstrained token-level prefix reuse in hybrid large language models. The key insight is that linear-attention mechanisms such as Gated DeltaNet can be viewed as a structured, lossy compression of the input prefix: gated recurrent updates progressively attenuate the contributions of earlier inputs. Consequently, the recurrent state of a matched prefix can be well approximated by replaying only a short, recent suffix of that prefix. Tail-Replay exploits this property by caching the exact full-attention key-value cache while omitting recurrent-state checkpoints. On a cache hit, it reconstructs the linear-attention states by replaying a short, recent suffix of the matched prefix. As a result, the reuse boundary is determined by the shared tokens rather than by recurrent-state checkpoints. We evaluate Tail-Replay on three Gated DeltaNet-based hybrid models using the LongBench and RULER benchmarks. With only a 5--10\% replay budget, it retains 92.8--99.9\% of full-prefill quality on LongBench and RULER. For serving efficiency, we evaluate time-to-first-token speedups across multiple matched-prefix lengths---8K, 16K, and 32K. The speedup grows with prefix length, reaching $9.1$--$14.3\times$ over full prefill at 32K.

↑