HyQuant:面向大语言模型注意力机制的混合精度量化方法
HyQuant: Hybrid-Precision Quantization for LLM Attention
- Xiamen University(厦门大学)
- Tencent Penglai Lab(腾讯蓬莱实验室)
- Shanghai Jiao Tong University(上海交通大学)
- Xi’an Jiaotong University(西安交通大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
HyQuant是一种面向LLM注意力的混合精度量化框架,通过保留关键区域高精度、量化其余部分,在各类任务上实现近乎无损精度,提升效率。
中文摘要 AI 辅助
量化技术已广泛应用于大语言模型(LLM)的训练与推理过程,以降低成本并提升效率。然而,注意力模块的低位宽量化在极低位宽下常引入较大误差,导致性能下降。现有方法主要依赖平滑技术处理异常值,而我们提出一种混合量化设计,以更好地平衡精度与效率。具体而言,我们提出HyQuant——一种面向LLM注意力的高效混合量化框架。HyQuant将大部分注意力状态量化为低位宽格式,同时保留一小部分垂直线(vertical-line)token和局部窗口状态的高精度表示。这些对精度至关重要的区域通过轻量级的垂直线感知注意力模式信号进行选择,以有限的开销降低量化误差。在Prefill阶段,HyQuant采用混合精度量化注意力算子,该算子以全精度保留垂直线token和局部滑动窗口,同时量化其余上下文。在Decode阶段,HyQuant将相同原理应用于KV缓存压缩,并将KV反量化与注意力计算融合,以提升内存与硬件效率。在各类任务、模型和数据集上,HyQuant以极其简洁的设计维持了近乎无损的精度,证明了混合量化应用于LLM注意力的效率与实际可行性。代码可获取于此:this https URL。
英文摘要
Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose \textbf{HyQuant}, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .