arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AlignQuant:面向高效大语言模型生成的瓦片对齐混合精度量化

AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generation

Hanzhi Zhang, Qiao Zhang, Qinglei Cao, Heng Fan, Yan Huang, Kewei Sha, Yunhe Feng

arXiv 2610.07457首次发表:更新:

发表机构

University of North Texas; Saint Louis University(北德克萨斯大学; 圣路易斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AlignQuant通过GPU兼容的二维权重瓦片统一精度分配、存储与执行,解决混合精度量化与GPU计算单元的边界冲突,实现高达2.50倍生成加速并保持模型质量。

AI 中文摘要

细粒度混合精度量化有望实现高效的大语言模型推理,但局部精度选择可能与GPU常规存储和计算单元冲突。这种精度边界不匹配限制了压缩向实际加速的转化。我们提出AlignQuant,一种训练后量化方法,采用GPU兼容的二维权重瓦片作为精度分配、紧凑存储和执行的通用单元。这种共享划分使得精度能够跟随输出通道内的敏感性。联合预填充/解码校准通过投影输出扰动(由量化激活下的语言模型损失梯度加权)来评估精度降低。相位归一化分数在模型级权重存储预算下,优先为对任一相位重要的瓦片分配更高精度。每个瓦片存储一种选定的表示,而相位专用内核复用打包模型,并将低位权重扩展为INT8计算,配合8位激活。在涵盖3B至14B参数的四个大语言模型上,AlignQuant相对于BF16实现了高达2.50倍的生成加速,同时保持模型质量。评估还覆盖了三种GPU和最长64K令牌的上下文。这些结果表明,局部精度灵活性和常规GPU执行可以通过共享瓦片单元共存。实现可在该https URL获取。

英文摘要

Fine-grained mixed-precision quantization promises efficient large language model inference, but local precision choices can conflict with regular GPU storage and computation units. This precision-boundary mismatch limits the translation of compression into practical acceleration. We introduce AlignQuant, a post-training quantization method that uses GPU-compatible two-dimensional weight tiles as the common unit of precision allocation, compact storage, and execution. This shared partition lets precision follow sensitivity within output channels. Joint prefill/decode calibration scores precision reductions using projection-output perturbations weighted by language-model loss gradients under quantized activations. Phase-normalized scores prioritize higher precision for tiles important to either phase under a model-wide weight-storage budget. Each tile stores one selected representation, while phase-specialized kernels reuse the packed model and expand lower-bit weights for INT8 computation with 8-bit activations. Across four LLMs spanning 3B to 14B parameters, AlignQuant achieves up to $2.50\times$ generation speedup over BF16 while preserving model quality. Evaluations further cover three GPUs and contexts up to 64K tokens. These results show that local precision flexibility and regular GPU execution can coexist through a shared tile unit. The implementation is available at https://github.com/HanzhiZhang-Ulrica/AlignQuant.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑