让线性状态遗忘遥远的过去:面向混合LLM的基于后缀重放的前缀缓存
Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs
浏览论文内容
中文总结 AI 辅助
针对混合LLM前缀缓存受限问题,提出SuffixReplay系统,通过重放近期后缀近似线性状态,无需检查点,实现任意边界前缀重用,在保持质量的同时显著降低存储和延迟。
中文摘要 AI 辅助
混合LLM将全注意力层与线性注意力层交错排列,以降低长上下文推理成本,但这种结构使前缀缓存变得复杂。全注意力KV缓存是令牌可寻址的,而线性注意力层维护的循环状态无法回滚到任意前缀边界。现有系统通过物化循环状态检查点,将前缀重用限制在检查点对齐的位置。我们提出了SuffixReplay,这是首个前缀缓存系统,它使混合LLM能够在每个缓存支持的页面边界重用缓存前缀,而无需物化循环状态检查点。我们的关键见解是让线性状态遗忘遥远的过去。现代线性注意力机制使用循环衰减和门控来减弱旧输入的影响。因此,SuffixReplay不是为每个前缀边界设置检查点,而是通过仅重放层输入隐藏状态的近期后缀(我们将其保留为锚点)来近似匹配边界处的状态。在算法层面,SuffixReplay结合了层级和令牌级锚点稀疏性以及有界重放预算,以控制存储、计算和质量。在系统层面,它使用独立管理的锚点sidecar和流水线重放路径,将锚点移动和状态重建与原生服务流水线重叠。我们在三个混合LLM上评估了SuffixReplay:OLMo-Hybrid-7B、Qwen3.5-4B和Qwen3.6-27B-FP8。在这些模型中,SuffixReplay在LongBench和RULER上平均保留了91.4%-100%的全预填充质量,同时仅使用SGLang默认8192令牌检查点缓存的0.36-0.51倍摊销每令牌存储。集成到SGLang后,SuffixReplay在分支工作负载上将中位TTFT降低了15%-70%,在工作集超过HBM时保持SGLang吞吐量的2.3-4.3倍,并在高命中续传流量上与SGLang匹配。
英文摘要
Hybrid LLMs interleave full-attention layers with linear-attention layers to reduce long-context inference cost, but this structure complicates prefix caching. Full-attention KV caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbitrary prefix boundaries. Existing systems materialize recurrent-state checkpoints, restricting prefix reuse to checkpoint-aligned positions. We present SuffixReplay, the first prefix caching system that lets hybrid LLMs reuse cached prefixes at every cache-supported page boundary without materializing recurrent-state checkpoints. Our key insight is to just let linear states forget the distant past. Modern linear-attention mechanisms use recurrent decay and gating to attenuate the influence of old inputs. Therefore, instead of checkpointing every prefix boundary, SuffixReplay approximates the state at a matched boundary by replaying only a recent suffix of the layer's input hidden states, which we retain as anchors. At the algorithmic level, SuffixReplay combines layer-wise and token-wise anchor sparsity with a bounded replay budget to control storage, computation, and quality. At the system level, it uses an independently managed anchor sidecar and a pipelined replay path to overlap anchor movement and state reconstruction with the native serving pipeline. We evaluate SuffixReplay on three hybrid LLMs: OLMo-Hybrid-7B, Qwen3.5-4B, and Qwen3.6-27B-FP8. Across these models, SuffixReplay retains 91.4-100% of full-prefill quality on average across LongBench and RULER, while using only 0.36-0.51x the amortized per-token storage of SGLang's default 8192-token checkpoint cache. Integrated into SGLang, SuffixReplay reduces median TTFT by 15-70% on branching workloads, sustains 2.3-4.3x SGLang's throughput when the working set exceeds HBM, and matches SGLang on high-hit continuation traffic.
发表机构
- Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信(TeleAI)人工智能研究所)
- Shanghai Jiao Tong University(上海交通大学)
- Tsinghua University(清华大学)
- Individual Researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。