arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LinearKV:混合大语言模型中与位置无关缓存仅需一个缓存状态

LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs

Yirui Liu, Ruoling Qi, Longwen Wang, Xuaner Wu, Jian Chen, Yuxin Jin, Jiawei Shao, Xuelong Li

arXiv 2608.11231首次发表:更新:

发表机构

Institute of Artificial Intelligence, China Telecom (TeleAI); Shanghai Jiao Tong University; University at Buffalo(中国电信人工智能研究院(TeleAI); 上海交通大学; 纽约州立大学布法罗分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LinearKV 是一种无需训练的混合大语言模型与位置无关缓存框架,仅需单个缓存状态初始化,在三类混合模型、三类选择器上均优于或等效于同期的 HYPIC 方案,可大幅提升 LLM 服务效率。

AI 中文摘要

大语言模型(LLM)服务正日益受益于与位置无关缓存(PIC)带来的加速。然而,现有的PIC方法是为全注意力模型设计的,其核心操作基于以 token 为索引的 KV 缓存:匹配可复用的 token 块、拼接这些块的 KV 条目,以及选择性地重新计算少量 token 以恢复跨块上下文。混合大语言模型打破了这些基本逻辑——它们用线性循环替换了大部分注意力层,仅暴露固定大小的状态,不存在可拼接或局部修复的以 token 为索引的 KV。这引出了一个自然问题:PIC 能否为混合模型带来益处,以及需要满足什么条件?我们提出 LinearKV,一个无需训练的混合 PIC 框架。其核心见解是「解耦初始化」:每个线性层将其 K 个匹配的局部状态映射为单个初始状态,而全注意力层则按以往方式拼接其 KV。因此,LinearKV 兼容现有的 PIC 方法,可直接复用其 token 选择与重新计算逻辑。在该框架下,我们发现「单个缓存状态」足以作为线性层的初始化器。代数上有严格依据的替代方案——将所有 K 个缓存状态组合为精确的全前缀状态(如同期工作 HYPIC 所做)——并无必要,且在部分架构上甚至会产生负面影响。我们在三个混合模型和三种 PIC 选择器上对比了这两种方案:在两个 GDN 模型上,两者表现相当,均恢复了大部分全量质量(最高达 92%);在 Mamba-2 模型上,精确组合在所有选择器下均失效——例如在 EPIC 选择器下,仅恢复 46.6% 的全量质量,而单个缓存块初始化器则恢复 86.8%。单个状态初始化器成本更低,将首 token 生成时间缩短至全预填充的 0.46 倍,而精确组合则额外增加 5%–17% 的开销;该结果在 8K–32K 长度的 LongBench 问答任务和 RULER 任务上均成立。

英文摘要

LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention models, where a token-indexed KV cache underlies its core operations: matching reusable token chunks, concatenating their KV entries, and selectively recomputing a few tokens to restore cross-chunk context. Hybrid LLMs break these primitives---they replace most attention layers with linear recurrences that expose only a fixed-size state, leaving no token-indexed KV to concatenate or to locally repair. This raises a natural question: can PIC benefit hybrid models, and what would it take? We present LinearKV, a training-free hybrid-PIC framework. Its key insight is a \emph{decoupled initialization}: each linear layer maps its $K$ matched local states to a single initial state, while full-attention layers concatenate their KV as before. LinearKV is therefore compatible with existing PIC methods, reusing their token selection and recomputation as-is. Under this framework, we find that a \emph{single cached state} suffices as the linear layer's initializer. The algebraically principled alternative---composing all $K$ cached states into the exact full-prefix state, as concurrent work HYPIC does---is unnecessary and, on some architectures, even harmful. We compare the two across three hybrid models and three PIC selectors. On the two GDN models the two tie, both recovering most of full quality (up to $92\%$); on the Mamba-2 model, exact composition instead collapses under every selector---under EPIC, for instance, it recovers only $46.6\%$ of full quality, versus $86.8\%$ for a single cached block initializer. A single state initializer is also cheaper, cutting time-to-first-token to $0.46\times$ full prefill versus a further $5$--$17\%$ overhead for exact composition; results hold across LongBench QA and RULER at 8K--32K.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑