arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03867cs.AR

面向高效低比特大语言模型推理的异构感知微缩放技术

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference

Junyi Luo, Xinting Jiang, Tai-Hao Wen, Ruichen Qi, Minxing Chu, Hongyi Wu, Gregory Kielian, Ben Laurie, Qirui Zhang, Quan Cheng, Dennis Sylvester, Mehdi Saligane

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对低比特LLM推理的MX格式精度损失问题,提出异构感知的AdaMX格式与加速器,在不增加EBW的前提下提升精度、降低存储能耗,在多类LLM及多模态模型上表现优异。

中文摘要 AI 辅助

微缩放(MX)目前是低比特大语言模型(LLM)推理的标准方案,其4比特形式MXFP4仍会损失大量精度,原因是现有MX格式要么在块间固定元素格式,要么固定精度恢复方案,仅能捕获有限的量化异构性。量化异构性体现在两个层面:1)块间层面,各块的优选元素格式与精度恢复方案存在差异;2)操作数层面,权重与激活函数需要不同的编码方式。我们提出AdaMX(自适应微缩放),一种异构感知的格式与加速器,它在等效比特宽度(EBW)不增加的前提下,为每个块选择精度恢复方案,为每个操作数选择表示形式。该设计覆盖两种块大小,提供更高精度的工作点与更低EBW的工作点,可节省存储。我们采用22nm FD-SOI工艺实现了AI加速器原型,包含所提出的解码器、计算单元与量化逻辑。与仅使用FP4乘法器的同配置MXFP4加速器相比,AdaMX仅增加约1%的系统能耗;在低EBW工作点下,AdaMX精度仍高于基线,同时降低了内存占用与能耗。在3B至70B规模的LLM上,AdaMX消除了MXFP4在常识任务上83%的精度损失、在MMLU任务上82%的精度损失,以及NVFP4损失的43%与27%。AdaMX还可推广至多模态模型,在Gemma-4 12B上,其在全部四个视觉-语言基准测试中均优于MXFP4,且保留了FP16精度的96%。

英文摘要

Microscaling (MX) is now the standard for low-bit large language model (LLM) inference. Its 4-bit form MXFP4 still loses substantial accuracy, because existing MX formats fix either the element format or the precision-recovery scheme across blocks, and thus capture only limited quantization heterogeneity. Quantization heterogeneity appears at two levels: 1) across blocks, the preferred element format and precision-recovery scheme vary; 2) across operands, weights and activations require different encoding. We introduce AdaMX (Adaptive Microscaling), a heterogeneity-aware format and accelerator. It selects the precision-recovery scheme per block and the representation per operand, at no increase in equivalent bit width (EBW). One design covers two block sizes, giving a higher-accuracy operating point and a lower-EBW operating point that saves storage. We implement a 22nm FD-SOI AI accelerator prototype with the proposed decoder, computing unit, and quantization logic. Against an otherwise identical MXFP4 accelerator with FP4-only multipliers, AdaMX adds about 1% system energy. At the lower-EBW point, AdaMX stays more accurate than the baseline while lowering both memory footprint and energy. Across LLMs from 3B to 70B, AdaMX removes 83% of the MXFP4 accuracy loss on commonsense and 82% on MMLU, and 43% and 27% of the NVFP4 loss. AdaMX also generalizes to multimodal models. On Gemma-4 12B, it leads MXFP4 on all four vision-language benchmarks and keeps up to 96% of FP16 accuracy.

↑