发表机构
Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出DUET模型,通过噪声级专家二重奏协调两步视频生成中质量与多样性的权衡,经改进的DUET+进一步提升质量并保留多样性优势,效果显著。
AI 中文摘要
扩散模型近年来已实现高质量的视频生成,但迭代采样的高成本阻碍了其实际部署。少步蒸馏可缓解该成本,但在其两大主流范式间存在质量-多样性的权衡:轨迹级蒸馏(如sCM)偏向多样性,而分布级蒸馏(如DMD)偏向质量。针对极端两步视频生成任务,我们提出DUET,通过噪声级的专家二重奏协调两大范式:sCM专家负责高噪声步骤以构建多样结构,DMD专家负责低噪声步骤以优化外观细节。由于两位专家以各自原生目标独立训练,DUET规避了损失级组合的优化难题,可同时实现质量与多样性而非权衡取舍。我们进一步识别出中继接口和高噪声阶段为剩余瓶颈,并用RL引导的专家适配解决,得到DUET+。以Wan2.1-T2V-1.3B为骨干,DUET将sCM的两步生成质量提升至接近DMD的水平,同时保留其几乎全部的结构多样性——约为DMD的两倍;DUET+则进一步提升整体质量,同时保留该多样性优势。综上,这些结果确立了噪声级专家专业化作为一种简单、有效的范式,用于协调两步视频生成中的多样性与质量。
英文摘要
Diffusion models have enabled high-quality video generation in recent years, but the high cost of iterative sampling hinders their practical deployment. Few-step distillation alleviates this cost, yet exposes a quality--diversity trade-off between its two dominant paradigms: trajectory-level distillation (e.g., sCM) favors diversity, whereas distribution-level distillation (e.g., DMD) favors quality. Targeting extreme two-step video generation, we introduce DUET, which reconciles the two paradigms through a noise-level duet of experts: an sCM expert takes the high-noise step to lay out diverse structure, and a DMD expert takes the low-noise step to refine appearance detail. Since the two experts are trained independently with their native objectives, DUET sidesteps the optimization difficulties of loss-level combinations and delivers quality and diversity jointly rather than trading one for the other. We further identify the relay interface and the high-noise stage as the remaining bottlenecks, and address them with RL-guided expert adaptation, yielding DUET+. With the Wan2.1-T2V-1.3B backbone, DUET lifts the two-step quality of sCM close to the level of DMD while retaining nearly all of its structural diversity---about twice that of DMD---and DUET+ further improves overall quality while preserving this diversity advantage. Together, these results establish noise-level expert specialization as a simple, effective paradigm for reconciling diversity and quality in two-step video generation.