发表机构
MMLab, CUHK; ByteDance Inc.; HKUST; CPII under InnoHK(香港中文大学多媒体实验室; 字节跳动有限公司; 香港科技大学; 香港创科旗下的人工智能与机器人创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DuoMatching框架,通过联合-边缘分布匹配及LatentBridge和潜变量变化采样,提升少步视频生成的视觉质量、构图与语义对齐,人工偏好率超80%。
AI 中文摘要
流式视频生成受益于分布匹配蒸馏(DMD),该方法将视频帧的联合分布与视频教师模型对真实视频分布的近似进行匹配。尽管这种联合匹配缓解了自回归展开过程中的漂移问题,但在视觉质量和语义对齐方面仍存在局限性。为解决这些局限,我们提出了DuoMatching,一种通过统一的联合-边缘公式来近似真实视频分布的分布匹配框架。在现有联合匹配公式的基础上,额外的边缘匹配目标提供了来自图像生成器的专用帧级监督,从而传递其互补的视觉和语义先验。为了在视频生成中应用这种帧级监督,我们引入了LatentBridge来解决视频学生模型与图像教师模型之间的潜在表示不匹配问题。潜变量变化采样(Latent Variation Sampling)进一步将这种帧级监督分配到不同的时间片段中,从而减少冗余。实验表明,DuoMatching在基本保持运动动态的同时,提升了视觉质量、构图和语义对齐。人工评估显示,与所有评估的基线相比,整体偏好率超过80%。项目页面可在该https URL访问。
英文摘要
Streaming video generation has benefited from distribution matching distillation (DMD), which matches the joint distribution of video frames to a video teacher's approximation of the real video distribution. Although this joint matching mitigates drift during autoregressive rollouts, limitations remain in visual quality and semantic alignment. To address these limitations, we propose DuoMatching, a distribution matching framework that approximates the real video distribution through a unified joint-marginal formulation. On top of existing joint matching formulations, the additional marginal matching objective provides dedicated frame-level supervision from an image generator, transferring complementary visual and semantic priors from it. To apply this frame-level supervision in video generation, we introduce LatentBridge to resolve the latent representation mismatch between the video student and the image teacher. Latent Variation Sampling further distributes such frame-level supervision across distinct temporal segments, reducing redundancy. Experiments demonstrate that DuoMatching improves visual quality, composition, and semantic alignment while largely preserving motion dynamics. Human evaluations show overall preference rates above 80% against all evaluated baselines. The project page is available at https://johnzhan2023.github.io/DuoMatching/.