arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KroQuant:用于扩散Transformer高效训练后量化的克罗内克结构块变换

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

Yann Bouquet, Alireza Khodamoradi, Kristof Denolf, Mathieu Salzmann

arXiv 2607.21446首次发表:更新:

发表机构

EPFL; Advanced Micro Devices, Inc.(洛桑联邦理工学院; 超威半导体公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对扩散Transformer训练后量化到W4A4质量严重下降问题,提出KroQuant方法,对激活值3个2元素块应用克罗内克结构可逆变换,存储参数少,运行快,结合离线LoRaQ权重校准,输出更接近FP参考且保持或提高图像质量。

AI 中文摘要

扩散Transformer(DiTs)训练后量化到W4A4会严重降低输出质量,因为进入每个线性层的激活值包含4位格式无法表示的异常值。标准修复方法是在量化激活值及其权重之前对激活值应用可逆线性变换及其逆变换。块之间的归一化层迫使此变换在每个去噪步骤在线运行,使其推理计算成本成为约束设计的关键因素。现有方法在量化质量和推理成本之间进行权衡。我们提出了KroQuant,一种PTQ方法,它对激活值的每个32元素块应用学习到的克罗内克结构可逆变换,存储的参数少于逐通道缩放的一半。该块局部结构以小张量核GEMM运行,在MI350 GPU上,KroQuant量化内核比SmoothQuant内核快14%。离线LoRaQ权重校准吸收剩余的逐权重量化误差。在W4A4(MXFP4e2)的PixArt-Σ、SANA和FLUX.1-schnell上,KroQuant在MJHQ-30K和SDCI上产生的输出比SVDQuant和LoRaQ更接近FP参考,同时保持或提高图像质量。

英文摘要

Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-bit formats cannot represent. The standard fix applies an invertible linear transform to the activations and its inverse to the weights before quantizing both. Normalization layers between blocks force this transform to run online at every denoising step, making its inference computation cost the binding design constraint. Existing options trade quantization quality for inference cost: per-channel scaling (SmoothQuant) is computationally cheap but impacts the magnitude of the channels, which can harm quantization accuracy; fixed Hadamard transforms yield better quantization accuracy but require large block sizes that incur a high online cost; learned full-$d$ invertible transforms calibrate best but entail a prohibitive dense $d \times d$ matrix multiplication (GEMM) per layer per step. We propose KroQuant, a PTQ method that applies a learned Kronecker-structured invertible transform to each 32-element block of the activation, storing less than half the parameters of per-channel scaling. The block-local structure runs as small tensor-core GEMMs, and on an MI350 GPU the KroQuant quantizer kernel is up to $14\%$ faster than the SmoothQuant kernel. Offline LoRaQ weight calibration then absorbs the residual per-weight quantization error. On PixArt-$Σ$, SANA, and FLUX.1-schnell at W4A4 (MXFP4e2), KroQuant produces outputs closer to the FP reference than SVDQuant and LoRaQ on MJHQ-30K and SDCI, while preserving or improving image quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑