arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33477cs.AIcs.PF

让线性状态遗忘遥远的过去:面向混合LLM的基于后缀重放的前缀缓存

Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs

Yirui Liu, Ruoling Qi, Xuaner Wu, Yuxin Jin, Jian Chen, Penghang Liu, Yafei Huang, Jiawei Shao, Xuelong Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对混合LLM前缀缓存受限问题,提出SuffixReplay系统,通过重放近期后缀近似线性状态,无需检查点,实现任意边界前缀重用,在保持质量的同时显著降低存储和延迟。

中文摘要 AI 辅助

混合LLM将全注意力层与线性注意力层交错排列,以降低长上下文推理成本,但这种结构使前缀缓存变得复杂。全注意力KV缓存是令牌可寻址的,而线性注意力层维护的循环状态无法回滚到任意前缀边界。现有系统通过物化循环状态检查点,将前缀重用限制在检查点对齐的位置。我们提出了SuffixReplay,这是首个前缀缓存系统,它使混合LLM能够在每个缓存支持的页面边界重用缓存前缀,而无需物化循环状态检查点。我们的关键见解是让线性状态遗忘遥远的过去。现代线性注意力机制使用循环衰减和门控来减弱旧输入的影响。因此,SuffixReplay不是为每个前缀边界设置检查点,而是通过仅重放层输入隐藏状态的近期后缀(我们将其保留为锚点)来近似匹配边界处的状态。在算法层面,SuffixReplay结合了层级和令牌级锚点稀疏性以及有界重放预算,以控制存储、计算和质量。在系统层面,它使用独立管理的锚点sidecar和流水线重放路径,将锚点移动和状态重建与原生服务流水线重叠。我们在三个混合LLM上评估了SuffixReplay:OLMo-Hybrid-7B、Qwen3.5-4B和Qwen3.6-27B-FP8。在这些模型中,SuffixReplay在LongBench和RULER上平均保留了91.4%-100%的全预填充质量,同时仅使用SGLang默认8192令牌检查点缓存的0.36-0.51倍摊销每令牌存储。集成到SGLang后,SuffixReplay在分支工作负载上将中位TTFT降低了15%-70%,在工作集超过HBM时保持SGLang吞吐量的2.3-4.3倍,并在高命中续传流量上与SGLang匹配。

英文摘要

Hybrid LLMs interleave full-attention layers with linear-attention layers to reduce long-context inference cost, but this structure complicates prefix caching. Full-attention KV caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbitrary prefix boundaries. Existing systems materialize recurrent-state checkpoints, restricting prefix reuse to checkpoint-aligned positions. We present SuffixReplay, the first prefix caching system that lets hybrid LLMs reuse cached prefixes at every cache-supported page boundary without materializing recurrent-state checkpoints. Our key insight is to just let linear states forget the distant past. Modern linear-attention mechanisms use recurrent decay and gating to attenuate the influence of old inputs. Therefore, instead of checkpointing every prefix boundary, SuffixReplay approximates the state at a matched boundary by replaying only a recent suffix of the layer's input hidden states, which we retain as anchors. At the algorithmic level, SuffixReplay combines layer-wise and token-wise anchor sparsity with a bounded replay budget to control storage, computation, and quality. At the system level, it uses an independently managed anchor sidecar and a pipelined replay path to overlap anchor movement and state reconstruction with the native serving pipeline. We evaluate SuffixReplay on three hybrid LLMs: OLMo-Hybrid-7B, Qwen3.5-4B, and Qwen3.6-27B-FP8. Across these models, SuffixReplay retains 91.4-100% of full-prefill quality on average across LongBench and RULER, while using only 0.36-0.51x the amortized per-token storage of SGLang's default 8192-token checkpoint cache. Integrated into SGLang, SuffixReplay reduces median TTFT by 15-70% on branching workloads, sustains 2.3-4.3x SGLang's throughput when the working set exceeds HBM, and matches SGLang on high-hit continuation traffic.

发表机构

  • Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信(TeleAI)人工智能研究所)
  • Shanghai Jiao Tong University(上海交通大学)
  • Tsinghua University(清华大学)
  • Individual Researcher(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

↑