VestigeKV:NoPE-MLA的KV缓存在退化分支中携带自身的逐出信号
VestigeKV: The NoPE-MLA KV Cache Carries Its Own Sparse-Attention Signal in a Vestigial Branch
浏览论文内容
中文总结 AI 辅助
VestigeKV利用NoPE-MLA的退化分支作为查询无关显著性信号,在不训练、量化或改变权重的情况下,高效压缩KV缓存,在长上下文下保持高检索准确率,且仅适用于NoPE架构。
中文摘要 AI 辅助
研究背景与问题:长生命周期的KV缓存必须在读取它的查询存在前被压缩;基于观察到的注意力(H2O、SnapKV)进行的选择在此处失效(NoPE MLA模型上的指针检索准确率为0.00-0.33),因为 token 的重要性尚未被观察到。提出的方法:针对Kimi Linear模型,VestigeKV通过查询无关信号(缓存本身携带的64维解耦分支,即RoPE的退化分支,NoPE训练将其重新用作显著性通道)进行逐出,读取每行的11%,将缓存分区:前m行保留在参与层,其余行精确移动(永不删除)到GPU驻留的存档,该存档可通过认证触发器每步访问;无需训练、量化,无权重或内核变更。实验设置与结果:无可测量开销:在8倍上下文扩展下检索保持1.00,32倍下从8k到65k上下文保持0.92,与全行列选择无差距;32倍时参与层占Kimi Linear每token缓存8.1KB中的0.25KB,存档保持位精确且GPU驻留,主机卸载是VRAM回收变体;召回层(标准配置)在128倍下保持1.00;Kimi K3据报道使用NoPE Gated-MLA变体,若其缓存布局匹配,该方法可合理扩展,本文仅针对已测量模型提出主张。NoPE专属特性:在RoPE MLA上使用相同算子会降至0.08(普通逐出为0.42);查询无关显著性仅在无旋转时存在(前1目标占token的2.3-6.7%,而旋转时为10.2-46.8%),且查询通用精确合并在RoPE下被证明不可能;所有阈值在数据前冻结,本文附带20个存档判决和8条封闭路径。
英文摘要
A long-lived KV cache must be compressed before the queries that will read it exist. Selection by observed attention collapses there: on a NoPE-MLA model, H2O and SnapKV retrieve 0.00 and 0.33 of needles at 8x compression, because a token's importance has not yet been observed. VestigeKV instead derives a sparse attention pattern from a signal the cache already carries, occupying the sparse-attention literature's one unoccupied quadrant: training-free and query-independent. In NoPE-MLA the 64-dimensional decoupled branch is a vestige of RoPE that training repurposes into a salience channel; reading 11% of each row, it partitions the cache into an attended tier and a GPU-resident archive that no row ever leaves, reachable each step by a certified, query-adaptive trigger. Nothing is trained and cache rows are never quantized, so every quality effect attributes to selection and scheduling. On Kimi Linear 48B, retrieval holds at 1.00 under 8x and 0.96 under 32x from 8k to 65k context, with zero gap to full-row selection, and the recall tier holds 128x at 1.00 (8k). Both tiers stay on the GPU, so the win is speed, not memory: the per-step scan reads ~26% of the bytes dense attention would, and on a two-node sglang deployment the crossover sits at ~40k context, reaching 1.18x at 256k and 1.39x at 496k. The mechanism is exclusive to NoPE: the identical operator on a RoPE MLA collapses to 0.08, query-independent salience exists only without rotation, and query-universal exact merging is provably impossible under RoPE. All thresholds were frozen before their data; 20 archived verdicts and 8 closed routes accompany the paper.
发表机构
- Yotta Labs
机构由 AI 辅助整理,请以论文原文为准。