两半胜于一:面向快速高质量图像生成的阶段式速度蒸馏
Two Halves are More than One: Phase-wise Velocity Distillation for Fast and High-Quality Image Generation
浏览论文内容
中文总结 AI 辅助
提出阶段式速度蒸馏(PVD),将生成时间线分为粗、细两阶段,各用半尺寸专家建模,在等效单次全骨干前向下提升图像生成质量,并显著降低参数与显存开销。
中文摘要 AI 辅助
近年来,基于扩散的图像生成骨干网络规模大幅增长,导致网络推理成本迅速增加。虽然扩散蒸馏技术可以减少推理步数,但在单个全骨干前向计算预算内生成高质量图像仍然具有挑战性。现有的一步法通常将这一预算分配给对整体式学生网络的单次评估。然而,用单个整体式映射来逼近异质的从粗到细的传输过程是困难的,且常常导致输出过度平滑。为解决这一问题,我们提出阶段式速度蒸馏(PVD),该方法将生成时间线划分为粗阶段和细阶段,并通过平均速度对每个阶段内的转换进行建模。为每个阶段分配一个专门的半尺寸专家,将结构组成与细节细化解耦,同时保持累积计算量等同于一次全骨干前向传播。我们证明,使用两个半尺寸的阶段专属专家优于单个全尺寸的整体式学生网络。在类条件图像生成任务上,PVD在ImageNet 256×256上实现了1.48的FID。在更复杂的文本到图像(T2I)任务中,PVD蒸馏模型(Stable Diffusion 3.5-Medium、FLUX.1-dev、Qwen-Image)产生的结果与多步教师模型相当,显著优于先前的蒸馏方法。此外,在评估的T2I骨干网络中,与相应教师模型相比,PVD将活跃参数减少了49.10%-50.89%,峰值显存减少了45.76%-48.36%。源代码和蒸馏模型可在该https URL获取。
英文摘要
Recent diffusion-based image generation backbones have grown substantially in scale, making the network inference cost increase rapidly. While diffusion distillation techniques can reduce the number of inference steps, high-quality image generation within a single full-backbone-forward compute budget remains challenging. Existing one-step methods typically allocate this budget to a single evaluation of a monolithic student. However, approximating the heterogeneous coarse-to-fine transport with a single monolithic mapping is difficult and often leads to over-smoothed outputs. To address this issue, we propose Phase-wise Velocity Distillation (PVD), which partitions the generation timeline into a coarse and a fine phase, and models the transition within each phase via the average velocity. A dedicated half-sized expert is assigned to each phase, decoupling structural composition from detail refinement while keeping the cumulative computation equivalent to one full-backbone forward pass. We show that the use of two half-sized phase-specific experts outperforms a single full-size monolithic student. On class-conditional image generation, PVD achieves an FID of 1.48 on ImageNet 256 x 256. On more complex text-to-image (T2I) tasks, PVD-distilled models (Stable Diffusion 3.5-Medium, FLUX.1-dev, Qwen-Image) produce results competitive with their multi-step teachers, significantly outperforming prior distillation methods. Moreover, across the evaluated T2I backbones, PVD reduces active parameters by 49.10-50.89% and peak VRAM by 45.76-48.36% compared to the corresponding teachers. Source code and distilled models are available at https://github.com/PolyU-VCLab/PVD.
发表机构
- The Hong Kong Polytechnic University(香港理工大学)
- OPPO Research Institute(OPPO研究院)
机构由 AI 辅助整理,请以论文原文为准。