发表机构
Robbyant; ZJU; HKUST; PolyU(睿冰科技(Robbyant); 浙江大学; 香港科技大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DSAQuant是面向VDMs的去噪阶段对齐量化感知训练框架,适配视频去噪阶段特性,在W3A3量化下将VBench平均分数提升最高6.60,优于SOTA基线。
AI 中文摘要
视频扩散模型(Video Diffusion Models, VDMs)在文本到视频生成领域已取得显著进展,但其高昂的内存与计算成本阻碍了实际部署。量化感知训练(Quantization-Aware Training, QAT)是一种有效方案,可在推理时无运行时开销地压缩和加速先进生成模型。然而,现有QAT方法在VDMs中面临独特挑战:虽常能保留提示语义、全局布局与粗略运动,但量化模型会严重降低视觉细节、纹理保真度与清晰度。本文将这种退化归因于传统量化流水线的时间步不敏感设计,其忽略了视频去噪的阶段性功能:VDMs中,早期去噪步骤主要构建全局结构与运动,而中晚期步骤则优化局部外观与高频细节。基于此,本文提出DSAQuant,一种面向VDMs的去噪阶段对齐量化感知训练框架。训练阶段,面向去噪阶段的监督在早期步骤保留教师蒸馏以实现稳定结构规划,同时将后期步骤转向目标驱动优化以增强细节重建;推理阶段,去噪门控引导在最终去噪步骤禁用分类器自由引导(Classifier-Free Guidance, CFG),防止其将量化诱导误差放大为高频伪影。在Wan与CogVideoX系列上,W4A4与W3A3设置下的大量实验表明,DSAQuant始终优于现有最优(SOTA)QAT基线,在激进的W3A3量化下将VBench平均分数提升最高达6.60,同时保持较强的文本-视频对齐能力。这些结果表明,有效的VDM量化不仅需要降低量化误差,还需使量化训练与推理适配视频扩散的阶段特性。
英文摘要
Video diffusion models (VDMs) have achieved impressive progress in text-to-video generation, but their high memory and computational costs hinder practical deployment. Quantization-aware training (QAT) is an effective solution for compressing and accelerating advanced generative models without runtime overhead at inference. However, existing QAT methods suffer from a distinctive challenge in VDMs: while they often preserve prompt semantics, global layout, and coarse motion, the quantized model severely degrades visual details, texture fidelity, and sharpness. In this paper, we trace this degradation to the timestep-agnostic design of conventional quantization pipelines, which overlooks the stage-wise functionality of video denoising. In VDMs, early denoising steps mainly establish global structure and motion, whereas middle and late steps refine local appearance and high-frequency details. Based on this insight, we propose DSAQuant, a Denoising-Stage-Aligned Quantization-aware training framework for VDMs. During training, Denoising-Stage Oriented Supervision preserves teacher distillation in early steps for stable structure planning, while shifting later steps toward target-driven optimization to enhance detail reconstruction. During inference, Denoising-Stage Gated Guidance disables CFG in the final denoising steps to prevent it from amplifying quantization-induced errors into high-frequency artifacts. Extensive experiments on the Wan and CogVideoX families under W4A4 and W3A3 settings show that DSAQuant consistently outperforms the SOTA QAT baseline, improving the VBench average score by up to 6.60 under aggressive W3A3 quantization while preserving strong text-video alignment. These results demonstrate that effective VDM quantization requires not only reducing quantization error, but also aligning quantization training and inference with the stage-wise nature of video diffusion.
CommentsProject page: \url{https://robbyant-research.github.io/DSAQuant/}; Code: \url{https://github.com/robbyant-research/DSAQuant}