arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11582cs.CV

OmniKVQuant:全模态大语言模型的KV缓存量化

OmniKVQuant: KV Cache Quantization for Omni-LLMs

Suho Yoo, Hyunjong Ok, Jongmin Choi, Jihoo Jung, Joon Son Chung

首次发表
浏览论文内容

中文总结 AI 辅助

针对全模态大模型KV缓存量化中的时间键漂移和异构值几何问题,提出无需训练的OmniKVQuant框架,通过窗口键量化和按模态值旋转,在Qwen2.5-Omni和Qwen3-Omni上实现2比特缓存并保持性能,附有高效解码内核。

中文摘要 AI 辅助

随着全模态大语言模型(Omni-LLMs)同时处理音频、视频和文本,其KV缓存的内存开销随之增长。KV缓存量化是文本专用大语言模型中的常用方法,但其在全模态大语言模型中的应用尚未被探索。本文分析了TurboQuant(一种基于旋转的KV缓存量化代表性方法)在多模态缓存上的表现,并发现了两个关键问题:时间键漂移和异构值几何。为解决这些问题,我们提出了OmniKVQuant,一个无需训练框架,其做法是:(i)在输入流的每个短窗口内设置键量化范围;(ii)按模态分别旋转值。在Qwen2.5-Omni和Qwen3-Omni上,OmniKVQuant实现了2比特KV缓存,同时在七个音视频基准上大幅保持了性能。我们进一步提供了一个融合的Triton解码内核,在注意力计算期间解包2比特缓存,因此无需构建密集的FP16缓存。代码:此https URL。

英文摘要

As Omni-modal large language models (Omni-LLMs) take in audio, video and text together, their KV cache memory cost grows. KV cache quantization is the de facto approach in text-only LLMs, but its application to Omni-LLMs remains unexplored. In this paper, we analyze how TurboQuant, a representative rotation-based KV cache quantization method, behaves on multimodal caches and identify two critical issues: temporal key drift and heterogeneous value geometry. To address these, we propose OmniKVQuant, a training-free framework that (i) sets the key quantization range over each short window of the input stream; and (ii) rotates values separately per modality. On Qwen2.5-Omni and Qwen3-Omni, OmniKVQuant enables 2-bit KV caches while substantially preserving performance across seven audio-visual benchmarks. We further provide a fused Triton decode kernel that unpacks the 2-bit cache during attention, so no dense FP16 cache is ever built. Code: https://github.com/kaistmm/OmniKVQuant

发表机构

  • KAIST(韩国科学技术院)
  • POSTECH(浦项科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑