arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于样本引导分布匹配的视频生成联合对齐与蒸馏

Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching

Jiuzhou Lin, Junlong Wu, Fei Zuo, Huan Ouyang, Dewen Fan, Boheng Zhang, Huaiqing Wang, Jia Sun, Fan Yang, Houde Liu, Kehai Chen, Min Zhang, Tingting Gao, Han Li

arXiv 2609.04283首次发表:更新:

发表机构

Tsinghua University; Kuaishou Technology; Harbin Institute of Technology (ShenZhen); Beijing University of Posts and Telecommunications(清华大学; 快手科技; 哈尔滨工业大学(深圳); 北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视频生成模型RL方法计算开销大、易致模型崩溃的问题,提出基于分布匹配的DM-Align单阶段框架,协同蒸馏与偏好引导梯度,提升生成质量与偏好对齐效果。

AI 中文摘要

将视频生成模型对齐人类偏好高度依赖强化学习(RL),但该方法存在计算开销巨大的问题。现有工作流程通常将RL与蒸馏视为不相关的阶段:在蒸馏前应用RL会产生过高的计算成本,而在蒸馏后应用RL则常导致模型崩溃。为克服这些局限,我们提出一种基于分布匹配(DM)的统一单阶段优化框架。在标准DM框架中,蒸馏通过最小化真实模型与生成模型之间差距的梯度方向更新模型,引导生成结果向清晰、高保真方向发展。在此基础上,我们引入DM-Align,该方法推导了互补的梯度方向,以引导模型向人类偏好样本靠拢。受DPO和GRPO启发,我们的方法利用分布差距(由偏好对或组内探索构建)直接构造此偏好引导梯度。通过协同这两种梯度方向,我们的方法无需传统RL中固有的多步奖励评估和复杂的ODE-SDE转换。在多个基础视频模型上进行的全面实验表明,这种样本引导框架可稳健提升蒸馏质量和偏好对齐效果,始终优于单独的变体方法及顺序两阶段流程。

英文摘要

Aligning video generative models to human preferences heavily relies on Reinforcement Learning (RL), which suffers from extensive computational overhead. Existing workflows typically treat RL and distillation as disconnected stages: applying RL before distillation incurs prohibitive computational costs, whereas applying RL after distillation frequently leads to model collapse. To overcome these limitations, we propose a unified, single-stage optimization framework grounded in Distribution Matching (DM). In the standard DM framework, distillation updates the model via a gradient direction that minimizes the gap between the real and fake models, guiding generations toward clarity and high fidelity. Building upon this, we introduce DM-Align, which derives a complementary gradient direction to guide the model toward human-preferred samples. Inspired by DPO and GRPO, our method leverages the distributional gap -- formulated from either preference pairs or intra-group exploration -- to directly construct this preference-guided gradient. By synergizing these two gradient directions, our approach eliminates the need for multi-step reward evaluation and complex ODE-SDE conversions inherent in traditional RL. Comprehensive experiments across multiple foundational video models demonstrate that this sample-guided framework robustly enhances both distillation quality and preference alignment, consistently outperforming both standalone variants and sequential two-stage pipelines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑