发表机构
Tsinghua University; Kuaishou Technology; Harbin Institute of Technology (ShenZhen); Beijing University of Posts and Telecommunications(清华大学; 快手科技; 哈尔滨工业大学(深圳); 北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对视频生成模型RL方法计算开销大、易致模型崩溃的问题,提出基于分布匹配的DM-Align单阶段框架,协同蒸馏与偏好引导梯度,提升生成质量与偏好对齐效果。
AI 中文摘要
将视频生成模型对齐人类偏好高度依赖强化学习(RL),但该方法存在计算开销巨大的问题。现有工作流程通常将RL与蒸馏视为不相关的阶段:在蒸馏前应用RL会产生过高的计算成本,而在蒸馏后应用RL则常导致模型崩溃。为克服这些局限,我们提出一种基于分布匹配(DM)的统一单阶段优化框架。在标准DM框架中,蒸馏通过最小化真实模型与生成模型之间差距的梯度方向更新模型,引导生成结果向清晰、高保真方向发展。在此基础上,我们引入DM-Align,该方法推导了互补的梯度方向,以引导模型向人类偏好样本靠拢。受DPO和GRPO启发,我们的方法利用分布差距(由偏好对或组内探索构建)直接构造此偏好引导梯度。通过协同这两种梯度方向,我们的方法无需传统RL中固有的多步奖励评估和复杂的ODE-SDE转换。在多个基础视频模型上进行的全面实验表明,这种样本引导框架可稳健提升蒸馏质量和偏好对齐效果,始终优于单独的变体方法及顺序两阶段流程。
英文摘要
Aligning video generative models to human preferences heavily relies on Reinforcement Learning (RL), which suffers from extensive computational overhead. Existing workflows typically treat RL and distillation as disconnected stages: applying RL before distillation incurs prohibitive computational costs, whereas applying RL after distillation frequently leads to model collapse. To overcome these limitations, we propose a unified, single-stage optimization framework grounded in Distribution Matching (DM). In the standard DM framework, distillation updates the model via a gradient direction that minimizes the gap between the real and fake models, guiding generations toward clarity and high fidelity. Building upon this, we introduce DM-Align, which derives a complementary gradient direction to guide the model toward human-preferred samples. Inspired by DPO and GRPO, our method leverages the distributional gap -- formulated from either preference pairs or intra-group exploration -- to directly construct this preference-guided gradient. By synergizing these two gradient directions, our approach eliminates the need for multi-step reward evaluation and complex ODE-SDE conversions inherent in traditional RL. Comprehensive experiments across multiple foundational video models demonstrate that this sample-guided framework robustly enhances both distillation quality and preference alignment, consistently outperforming both standalone variants and sequential two-stage pipelines.