AI 中文总结
PrismQuant通过量化器感知的旋转对齐激活特征与分组INT4子空间,以闭式最优解实现低比特量化,在Llama等模型上达到最先进性能并显著加速。
AI 中文摘要
较小的激活异常值并不一定意味着更好的低比特量化:它们与量化器的对齐方式才是关键。我们提出了PrismQuant,一种量化器感知的旋转框架,它将主导激活特征空间与非对称分组INT4的恒定组子空间对齐。仿射偏移量表示该子空间中的能量,而不会扩大组内的范围。我们将旋转设计表述为Ky Fan迹最大化问题,并推导出一个闭式解,该解对于这一对齐目标而言是可证明最优的。紧凑的Householder变换及其紧凑WY表示使得在可折叠和在线位置都能进行无梯度的构建和高效应用。一个预测性的范围定律进一步将未对齐的激活能量和组大小与量化相关的变异联系起来。在Llama、Qwen和Mistral上的实验涵盖了高达70B参数的密集模型和一个30B参数的混合专家模型。在W4A4KV4下,PrismQuant在Llama-3.2-3B上以困惑度和准确率在比较方法中达到了最先进水平。在Llama-3.1-70B上,它达到了3.85的困惑度和72.46%的平均零样本准确率,仅比全精度低0.22个百分点。在Llama-3.1-8B的部署研究中,我们优化的实现相对于匹配的FP16基线实现了1.51倍的预填充和1.22倍的CUDA图解码加速,解码峰值内存降低了56.34%,并且相对于Hadamard仅增加了2.35%的图解码延迟。代码可在该https URL获取。
英文摘要
Smaller activation outliers do not necessarily imply better low-bit quantization: their alignment with the quantizer matters. We introduce PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the constant group subspace of asymmetric grouped INT4. The affine offsets represent the energy in this subspace without widening the range within the group. We formulate rotation design as a Ky Fan trace maximization and derive a closed-form solution that is provably optimal for this alignment objective. Compact Householder transformations and their compact-WY representation enable gradient-free construction and efficient application at both foldable and online sites. A predictive range law further connects unaligned activation energy and group size to quantization-relevant variation. Experiments on Llama, Qwen, and Mistral span dense models up to 70B parameters and a 30B mixture-of-experts model. Under W4A4KV4, PrismQuant sets the state of the art on Llama-3.2-3B among the compared methods in both perplexity and accuracy. On Llama-3.1-70B, it attains 3.85 perplexity and 72.46% average zero-shot accuracy, only 0.22 percentage points below full precision. In the deployment study on Llama-3.1-8B, our optimized implementation achieves 1.51x prefill and 1.22x CUDA Graph decode speedups over matched FP16 baselines, with 56.34% lower decode peak memory and only 2.35% additional Graph decode latency over Hadamard. Code is available at https://github.com/ForeverBlue816/PrismQuant.
CommentsThe paper is currently under review. Code, checkpoints, and implementation details are available at: https://github.com/ForeverBlue816/PrismQuant