arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向INT2 KV缓存量化的输出感知旋转

Output-Aware Rotation for INT2 KV-Cache Quantization

Vincent-Daniel Yun, Woosang Lim, Minsoo Cheong, Sunwoo Lee, Murali Annavaram, Sai Praneeth Karimireddy, Sungjoo Yoo

arXiv 2608.02691首次发表:更新:

发表机构

University of Southern California; Seoul National University; Inha University(南加州大学; 首尔大学; 仁荷大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长上下文大模型KV缓存的INT2量化瓶颈,提出OptR输出感知旋转方法,通过分解误差、学习正交修正等,提升QuaRot等性能并增强长上下文检索能力,且开销可忽略。

AI 中文摘要

键值(KV)缓存已成为长上下文大语言模型推理中主要的内存和带宽瓶颈,这使得超低比特量化愈发重要。然而,现有的基于旋转的INT2方法在完整注意力读出之前优化缓存统计量或代理误差,而模型最终会受到注意力传播的误差以及输出投影W_O的影响。为解决这种不匹配,我们提出OptR,一种输出感知旋转方法,用于最小化W_O后的注意力-输出误差。OptR将W_O后的注意力-输出误差分解为键和值诱导的项,并通过完整的INT2量化和注意力路径学习每头的正交修正。OptR还应用注意力等效的键重参数化,以在不改变softmax分布的情况下减少大的通道偏移。在三个模型和五个推理与编码基准上,OptR始终提升QuaRot和OSCAR的性能,增强长上下文检索能力,同时保留分页KV缓存格式且推理开销可忽略不计。

英文摘要

The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization increasingly important. However, existing rotation-based INT2 methods optimize cache statistics or proxy errors before the complete attention readout, even though the model is ultimately affected by the error propagated through attention and the output projection $W_O$. To address this mismatch, we propose \textit{OptR}, an output-aware rotation method that minimizes post-$W_O$ attention-output error. OptR decomposes the post-$W_O$ attention-output error into key- and value-induced terms and learns per-head orthogonal corrections through the full INT2 quantization and attention path. OptR further applies an attention-equivalent key reparameterization to reduce large channel-wise offsets without changing the softmax distribution. Across three models and five reasoning and coding benchmarks, OptR consistently improves both QuaRot and OSCAR and strengthens long-context retrieval, while preserving the paged KV-cache format with negligible inference overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑