arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HiKV:用于大语言模型解码的具有硬件加速的分层重要性感知键值缓存

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding

Chao Fang, Jun Yin, Man Shi, Marian Verhelst

arXiv 2607.22389首次发表:更新:

发表机构

ESAT-MICAS, KU Leuven(KU莱顿大学ESAT-MICAS)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长上下文大语言模型解码时键值缓存的内存瓶颈,HiKV提出算法-硬件协同设计,通过分层重要性感知在两粒度压缩缓存,开发专用加速器统一两阶段加速,在大语言模型上评估实现加速、降能及减少内存访问,优于现有方法。

AI 中文摘要

随着长上下文大语言模型的迅速普及,解码过程中不断增长的键值缓存成为关键的内存瓶颈。为应对这一挑战,我们提出了HiKV,一种新颖的算法-硬件协同设计,通过分层重要性感知利用键值缓存冗余。算法上,HiKV在两个粒度上压缩键值缓存:第一阶段在固定预算内逐出不重要的令牌,第二阶段仅加载每个保留令牌的重要元素,实现单粒度无法达到的压缩率。架构上,我们开发了一个以可重构重要性排序器为中心的专用加速器,在每个阶段所需的不同排序数据路径之间切换,以最小开销在一个电路中统一两阶段加速。在代表性大语言模型上评估,HiKV在注意力计算中实现高达7.95倍的加速和90%的能量降低,精度损失可忽略不计仅1%。在等精度约束下,HiKV通过减少外部内存访问额外实现1.82至4.87倍的降低,优于基于重要性的现有方法。这些优势由仅增加8%系统面积的专用硬件组件实现。

英文摘要

With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck. To tackle this challenge, we propose HiKV, a novel algorithm-hardware co-design that exploits KV cache redundancy through hierarchical importance awareness. Algorithmically, HiKV compresses the KV cache at two granularities: Stage I evicts unimportant tokens within a fixed budget, and Stage II further loads only the significant elements of each retained token, reaching compression ratios unattainable at a single granularity. Architecturally, we develop a dedicated accelerator centered on a reconfigurable importance sorter that switches between the distinct sorting datapaths each stage requires, unifying the two-stage acceleration in one circuit with minimal overhead. Evaluated on representative LLMs, HiKV achieves up to 7.95x speedup and 90% energy reduction in the attention computation over the vanilla KV cache baseline within negligible 1% accuracy loss. Under iso-accuracy constraints, HiKV outperforms state-of-the-art importance-based methods by achieving an additional 1.82~4.87x reduction in external memory accesses. These benefits are enabled by specialized hardware components that add only 8% to the system area.

CommentsTo appear in the IEEE Transactions on Circuits and Systems I: Regular Papers (TCAS-I)

DOI:10.1109/TCSI.2026.3716602

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑