AI 中文总结
FlashQuant是面向异常值感知W4A16解码的内容共享执行框架,通过融合GPU内核实现稀疏-密集路径的片上复用,在内存受限的LLM解码中大幅降低异常值处理开销,获得显著加速比。
AI 中文摘要
低比特量化可减少大语言模型(LLM)推理的内存占用与计算成本,但幅度较大的异常值权重会引发显著的量化误差,降低模型精度。异常值感知量化通过将异常值保留为高精度、其余权重量化的方式解决该问题,进而生成低比特密集通用矩阵乘法(GEMM)路径与高精度稀疏稀疏矩阵乘法(SpMM)路径。现有实现虽共享激活值与输出,却在独立的GPU内核中执行上述路径,错失算子内复用机会并产生冗余的全局内存访问,在内存受限的解码工作负载中该低效性尤为突出。本文提出FlashQuant,一种面向异常值感知W4A16解码的内容共享执行框架,将密集GEMM与稀疏异常值SpMM路径融合为单个GPU内核,实现异构计算中激活值与输出块的片上复用。它引入三项关键技术:稀疏-密集分块,使异常值处理与密集GEMM分块对齐;Tile-COO异常值编码,支持高效稀疏访问并减少共享内存体冲突;流水线调度,将计算与数据移动重叠。实验表明,FlashQuant降低了异常值处理开销,相较于cuBLAS BF16实现了2.74倍至4.18倍的加速,相较于最强的未融合异常值感知基线实现了最高1.53倍的加速。
英文摘要
Low-bit quantization reduces the memory footprint and computational cost of large language model (LLM) inference. However, high-magnitude outlier weights can induce substantial quantization errors and degrade model accuracy. Outlier-aware quantization addresses this issue by retaining outliers in high precision while quantizing the remaining weights, resulting in a low-bit dense GEMM path and a high-precision sparse SpMM path. Existing implementations execute these paths in separate GPU kernels, despite their shared activations and outputs, thereby missing opportunities for intra-operator reuse and incurring redundant global-memory accesses. This inefficiency is particularly pronounced in memory-bound decoding workloads. We propose FlashQuant, a content-sharing execution framework for outlier-aware W4A16 decoding. FlashQuant fuses the dense GEMM and sparse outlier SpMM paths into a single GPU kernel, enabling on-chip reuse of activation and output tiles across heterogeneous computations. It introduces three key techniques: sparse-dense tiling, which aligns outlier processing with dense GEMM tiles; Tile-COO outlier encoding, which enables efficient sparse access and reduces shared-memory bank conflicts; and pipelined scheduling, which overlaps computation with data movement. Experiments show that FlashQuant reduces outlier-processing overhead, achieving $2.74\times - 4.18\times$ speedup over cuBLAS BF16 and up to $1.53\times$ speedup over the strongest unfused outlier-aware baseline.