发表机构
The University of Hong Kong; Huawei Technologies Co., Ltd.(香港大学; 华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多模态扩散Transformer的W4A4量化局限,提出统一框架\textit{方法},通过降低海森矩阵与激活噪声代理优化,以秩-4实现优于SVDQuant的量化性能,且成本更低、速度更快。
AI 中文摘要
在扩散Transformer中,低秩分支可通过将每个权重分解为低比特残差和高精度低秩分量,缓解4比特权重-激活(W4A4)后训练量化(PTQ)的损失。然而,现有的低秩PTQ方法要么分别优化低秩补偿和残差量化,通常需要更高的秩;要么依赖二阶权重更新,未显式建模激活量化误差,该误差在4比特量化下尤为明显。为解决这些局限,我们提出\textit{方法}(注:原文未给出具体方法名,保留\textit{方法}占位),这是一个将低秩辅助W4A4 PTQ建模为耦合校准问题的统一框架,并从联合目标中推导基于优化的求解器。消除输出侧低秩因子可得到\textit{降低的海森矩阵},其会折扣已被低秩分量捕获的残差误差,同时引入激活噪声代理以抑制激活量化误差。在五个扩散骨干网络上,秩-4 \textit{方法}在PSNR和LPIPS指标上始终优于秩-4 SVDQuant;在SANA-1.6B、FLUX.1-schnell和FLUX.1-dev上,其以8倍更小的秩、最高6.25倍更快的量化速度,超越了秩-32 SVDQuant。此外,在Qwen3-8B大语言模型(LLM)上,秩-4 \textit{方法}将MMLU准确率从61.50%提升至68.17%,优于秩-32 SVDQuant。总体而言,\textit{方法}以显著更低的秩和量化成本实现了更优的W4A4性能。
英文摘要
In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component. Existing low-rank PTQ approaches, however, either optimize low-rank compensation and residual quantization separately, often requiring higher ranks, or rely on second-order weight updates without explicitly modeling activation quantization error, which becomes particularly pronounced under 4-bit quantization. To address these limitations, we present \method{}, a unified framework modeling low-rank-assisted W4A4 PTQ as a coupled calibration problem and deriving optimization-based solvers from the joint objective. Eliminating the output-side low-rank factor yields a \emph{deflated Hessian} that discounts residual errors already captured by the low-rank component, while an activation-noise surrogate is incorporated to suppress activation quantization error. Across five diffusion backbones, rank-4 \method{} consistently outperforms rank-4 SVDQuant in PSNR and LPIPS. It further surpasses rank-32 SVDQuant on SANA-1.6B, FLUX.1-schnell, and FLUX.1-dev with an $8\times$ smaller rank and up to $6.25\times$ faster quantization. Furthermore, on the Qwen3-8B LLM, rank-4 \method{} improves MMLU accuracy from 61.50\% to 68.17\% over rank-32 SVDQuant. Overall, \method{} achieves better W4A4 performance with substantially lower rank and quantization cost.