从静态到动态:从图像到视频扩散模型的策略内蒸馏
From Static to Dynamic: On-Policy Distillation from Image to Video Diffusion Models
查看机构详情
- The University of Hong Kong(香港大学)
- Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出MILD框架,通过可学习线性连接器对齐图像与视频潜在空间,结合光流运动奖励,将图像专家能力蒸馏到视频扩散模型,在保持动态的同时提升生成质量与一致性。
中文摘要 AI 辅助
策略内蒸馏(OPD)通过沿学生自身生成轨迹的教师监督来专门化预训练的视频扩散模型。尽管大型视频模型是天然的教师,但开发专门的视频专家可能需要昂贵的视频数据和训练,而查询它们的延迟远高于查询图像专家。图像专家更易获得且查询成本更低,提供了一种成本效益高的替代方案,尤其适用于美学和OCR等主要与时间无关的能力,这些能力允许帧级监督。然而,异构的图像和视频潜在空间阻碍了对学生中间状态的直接监督,而图像专家缺乏跨帧运动监督,使得时间一致性容易受到帧级改进的影响。在本文中,我们提出了MILD,一种运动保持的图像到视频潜在蒸馏框架,它在保持预训练视频动态的同时转移专门的图像专业知识。MILD使用一个可学习的线性连接器,将学生的潜在状态和预测更新与图像专家的对齐,从而实现跨异构潜在空间的监督转移。我们进一步将图像引导的修正约束在预训练学生预测周围,以保持视频动态,并引入基于光流的运动奖励来提高运动质量和时间一致性。在专门的图像专家和多个视频学生骨干网络上,我们的方法始终优于视频教师OPD基线,进一步的研究证明了在连接器设计和异构架构之间的有效转移。这些结果确立了图像到视频蒸馏作为通过利用图像生成生态系统的多样且不断发展的能力来改进视频生成的有效途径。
英文摘要
On-policy distillation (OPD) specializes pretrained video diffusion models through teacher supervision along the student's own generation trajectory. Although large video models are natural teachers, developing specialized video experts can require costly video data and training, while querying them incurs substantially higher latency than querying image experts. More readily available and cheaper to query, image experts offer a cost-effective alternative, particularly for largely temporal-agnostic capabilities such as aesthetics and OCR that admit frame-level supervision. However, heterogeneous image and video latent spaces prevent direct supervision of intermediate student states, while image experts lack cross-frame motion supervision, making temporal consistency vulnerable to frame-level improvements. In this paper, we propose MILD, a Motion-Preserving Image-to-Video Latent Distillation framework that transfers specialized image expertise while preserving pretrained video dynamics. MILD uses a learnable linear connector that aligns student latent states and predicted updates with those of image experts, enabling supervision transfer across heterogeneous latent spaces. We further constrain image-guided corrections around the pretrained student's predictions to preserve video dynamics and incorporate an optical-flow-based motion reward to improve motion quality and temporal consistency. Across specialized image experts and multiple video-student backbones, our method consistently outperforms video-teacher OPD baselines, with further studies demonstrating effective transfer across connector designs and heterogeneous architectures. These results establish image-to-video distillation as an effective route to improving video generation by drawing on the diverse and evolving capabilities of the image-generation ecosystem.