发表机构
EPFL; Tsinghua University; MangoBoost Inc.; Korea University; Google DeepMind(洛桑联邦理工学院; 清华大学; 芒果助推公司; 韩国大学; 谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型推理中4位量化因异常值致精度降的问题,提出MXSens方法,基于列和层灵敏度分配混合尾数比特宽度,无需训练,利用MXINT块结构,在多模型任务中优于现有方法,平衡了量化的准确性与资源效率。
AI 中文摘要
4位量化可实现高效的大语言模型推理,但由于异常值会导致显著的精度下降。先前的工作通过数据旋转或混合精度整数量化来解决此问题,但通常依赖软件管理的缩放和频繁的反量化,带来大量开销。微缩放格式(如MXINT)通过在硬件中编码缩放来消除这些低效率,但仍与基于旋转的方法不兼容。我们的分析表明,异常值的严重程度各不相同,量化灵敏度在各层和各列中分布不均。这些见解促使我们采用一种细粒度、灵敏度引导的方法。我们引入了MXSens,这是一种无需训练的方法,它根据列和层的灵敏度分配混合尾数比特宽度(4/6/8),自然地利用了MXINT的块结构。MXSens在一系列模型和任务上优于现有量化方法。在W4A4KV4设置下,MXSens在LLaMA-2-70B和LLaMA-3-8B上分别实现了3.77和7.63的困惑度,在WikiText-2上比现有基线有显著改进。我们的工作在大语言模型量化的准确性和资源效率之间建立了新的平衡。
英文摘要
4-bit quantization enables efficient LLM inference, but suffers from significant accuracy degradation due to outliers. Prior work addresses this problem via data rotation or mixed-precision integer quantization, but often relies on software-managed scaling and frequent dequantization, incurring substantial overhead. Microscaling formats, such as MXINT, eliminate these inefficiencies by encoding scales in hardware, yet remain incompatible with rotation-based methods. Our analysis reveals that outliers vary in severity, from rare extremes to frequent mild deviations, and that quantization sensitivity is unevenly distributed across layers and columns. These insights motivate a fine-grained, sensitivity-guided approach. We introduce MXSens, a training-free method that assigns mixed mantissa bitwidths (4/6/8) based on column- and layer-wise sensitivity, naturally leveraging the block-wise structure of MXINT. MXSens outperforms state-of-the-art quantization methods across a range of models and tasks. Under the W4A4KV4 setting, MXSens achieves perplexities of 3.77 and 7.63 on LLaMA-2-70B and LLaMA-3-8B, respectively, substantially improving over existing baselines on WikiText-2. Our work establishes a new balance between accuracy and resource efficiency for LLM quantization.