OmniKVQuant:全模态大语言模型的KV缓存量化
OmniKVQuant: KV Cache Quantization for Omni-LLMs
浏览论文内容
中文总结 AI 辅助
针对全模态大模型KV缓存量化中的时间键漂移和异构值几何问题,提出无需训练的OmniKVQuant框架,通过窗口键量化和按模态值旋转,在Qwen2.5-Omni和Qwen3-Omni上实现2比特缓存并保持性能,附有高效解码内核。
中文摘要 AI 辅助
随着全模态大语言模型(Omni-LLMs)同时处理音频、视频和文本,其KV缓存的内存开销随之增长。KV缓存量化是文本专用大语言模型中的常用方法,但其在全模态大语言模型中的应用尚未被探索。本文分析了TurboQuant(一种基于旋转的KV缓存量化代表性方法)在多模态缓存上的表现,并发现了两个关键问题:时间键漂移和异构值几何。为解决这些问题,我们提出了OmniKVQuant,一个无需训练框架,其做法是:(i)在输入流的每个短窗口内设置键量化范围;(ii)按模态分别旋转值。在Qwen2.5-Omni和Qwen3-Omni上,OmniKVQuant实现了2比特KV缓存,同时在七个音视频基准上大幅保持了性能。我们进一步提供了一个融合的Triton解码内核,在注意力计算期间解包2比特缓存,因此无需构建密集的FP16缓存。代码:此https URL。
英文摘要
As Omni-modal large language models (Omni-LLMs) take in audio, video and text together, their KV cache memory cost grows. KV cache quantization is the de facto approach in text-only LLMs, but its application to Omni-LLMs remains unexplored. In this paper, we analyze how TurboQuant, a representative rotation-based KV cache quantization method, behaves on multimodal caches and identify two critical issues: temporal key drift and heterogeneous value geometry. To address these, we propose OmniKVQuant, a training-free framework that (i) sets the key quantization range over each short window of the input stream; and (ii) rotates values separately per modality. On Qwen2.5-Omni and Qwen3-Omni, OmniKVQuant enables 2-bit KV caches while substantially preserving performance across seven audio-visual benchmarks. We further provide a fused Triton decode kernel that unpacks the 2-bit cache during attention, so no dense FP16 cache is ever built. Code: https://github.com/kaistmm/OmniKVQuant
发表机构
- KAIST(韩国科学技术院)
- POSTECH(浦项科技大学)
机构由 AI 辅助整理,请以论文原文为准。