EFQ-Softmax:用于Softmax的无指数量化
EFQ-Softmax: Exp-Free Quantization for Softmax
- Huawei Technologies Co., Ltd(华为技术有限公司)
- Shandong University of Finance and Economics(山东财经大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
EFQ-Softmax提出无指数低位概率生成方法,直接映射注意力分数为E2M1操作数,替代先指数后量化路径,在保持模型质量的同时降低延迟。
AI中文摘要:
低位注意力通过将$QK^\top$和$PV$矩阵乘法移至FP8或FP4矩阵引擎来加速Transformer推理。然而,softmax路径通常以更高精度计算偏移分数的指数,形成临时概率块,并在低位$PV$乘法之前对其进行量化。这种先指数后量化的路径在高精度概率生成器与低位矩阵消费者之间造成了不匹配。我们提出EFQ-Softmax(用于Softmax的无指数量化),一种低位概率生成方法,直接将偏移注意力分数映射为块缩放E2M1操作数。对于每个微缩放块,EFQ-Softmax从局部最大值中选择仅指数缩放,将偏移分数映射到归一化残差域,并使用单一仿射规则生成非负E2M1概率码。所得操作数在$\widetilde{P}V$分子更新和$\widetilde{P}\mathbf{1}$分母更新中一致使用。FlashAttention风格的行最大值更新、历史重新缩放、高精度累加和最终归一化保持不变。我们在Qwen3-8B、Qwen3-VL-8B-Instruct和WAN2.2-TI2V-5B上评估端到端质量,并单独在A5向量单元上测量内核级性能。EFQ-Softmax将Qwen3-8B七任务平均值从MXFP4的0.6749提升至0.6773,将Qwen3-VL九任务平均值从0.7826提升至0.8000。在WAN2.2上,它在VBench下保持了与FP16和MXFP4基线相当的时间一致性和视觉质量。在A5向量单元上,EFQ-Softmax在16K到128K序列长度上将融合概率生成内核的向量阶段延迟平均降低40.33%。这些结果表明,直接低位概率生成可以取代传统的先指数后量化路径,同时保持端到端模型质量。
英文摘要:
Low-bit attention accelerates Transformer inference by moving the $QK^\top$ and $PV$ matrix multiplications to FP8 or FP4 matrix engines. However, the softmax path often evaluates shifted-score exponentials in higher precision, forms a temporary probability block, and quantizes it before low-bit $PV$ multiplication. This exp-then-quantize path creates a mismatch between a high-precision probability producer and a low-bit matrix consumer. We propose EFQ-Softmax (Exp-Free Quantization for Softmax), a low-bit probability-generation method that directly maps shifted attention scores to block-scaled E2M1 operands. For each microscaling block, EFQ-Softmax selects an exponent-only scale from the local maximum, maps the shifted scores to a normalized residual domain, and generates nonnegative E2M1 probability codes using a single affine rule. The resulting operand is used consistently in both the $\widetilde{P}V$ numerator update and the $\widetilde{P}\mathbf{1}$ denominator update. The FlashAttention-style row-maximum update, historical rescaling, high-precision accumulation, and final normalization remain unchanged. We evaluate end-to-end quality on Qwen3-8B, Qwen3-VL-8B-Instruct, and WAN2.2-TI2V-5B, and separately measure kernel-level performance on the A5 vector unit. EFQ-Softmax improves the Qwen3-8B seven-task mean from 0.6749 with MXFP4 to 0.6773 and the Qwen3-VL nine-task mean from 0.7826 to 0.8000. On WAN2.2, it maintains temporal consistency and visual quality comparable to the FP16 and MXFP4 baselines under VBench. On the A5 vector unit, EFQ-Softmax reduces the vector-stage latency of the fused probability-generation kernel by 40.33% on average across sequence lengths from 16K to 128K. These results show that direct low-bit probability generation can replace the conventional exp-then-quantize path while preserving end-to-end model quality.