发表机构
Arizona State University; Twitch; Stanford University; eBay; NewsBreak; Microsoft; Columbia University; University of Southern California; Carnegie Mellon University(亚利桑那州立大学; Twitch; 斯坦福大学; eBay; NewsBreak; 微软; 哥伦比亚大学; 南加州大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文首次系统综述视频生成模型的后训练与对齐方法,提出统一框架,将方法分为四类,并讨论数据集、评估及开放挑战,以提升可控性与可靠性。
AI 中文摘要
视频生成已从短时、低质量的片段迅速发展到具有复杂时空动力学的高分辨率、长时长序列。尽管通过大规模预训练学习到了强大的生成先验,预训练视频模型往往无法可靠地遵循人类意图、保持时间连贯性或满足物理与安全约束。与图像和文本生成相比,视频生成中的对齐面临独特挑战,包括随时间累积的误差、运动-外观耦合、多目标权衡以及时间属性的监督有限。这些挑战促使了系统性的后训练策略,在不从头重新训练的情况下调整预训练模型。在本综述中,我们首次对视频生成模型中的后训练和对齐进行了全面回顾。我们将后训练构建为一个统一框架,并根据对齐信号的实施方式区分隐式对齐和显式对齐。基于此视角,我们将现有方法组织为四大类:监督微调方法、自训练与蒸馏方法、基于偏好和奖励的方法以及推理时方法。该分类法提供了对齐信号如何在训练和部署中塑造模型行为的连贯视图。除方法论进展外,我们回顾了常用的数据集、基准和评估实践,并讨论了开放挑战,如可扩展奖励设计、长时程时间一致性、稳定性-表现力权衡以及安全感知生成。本综述旨在为推进可控且可靠的视频生成模型提供结构化的概念基础和实用指导。
英文摘要
Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, maintain temporal coherence, or satisfy physical and safety constraints. Compared with image and text generation, alignment in video generation presents unique challenges, including error accumulation over time, motion-appearance coupling, multi-objective trade-offs, and limited supervision for temporal properties. These challenges motivate systematic post-training strategies that adapt pretrained models without retraining them from scratch. In this survey, we present the first comprehensive review of post-training and alignment in video generation models. We frame post-training as a unifying framework and distinguish between implicit alignment and explicit alignment based on how alignment signals are enforced. From this perspective, we organize existing approaches into four broad categories: supervised fine-tuning methods, self-training and distillation methods, preference- and reward-based methods, and inference-time methods. This taxonomy provides a coherent view of how alignment signals shape model behavior across both training and deployment. Beyond methodological advances, we review commonly used datasets, benchmarks, and evaluation practices, and discuss open challenges such as scalable reward design, long-horizon temporal consistency, stability-expressiveness trade-offs, and safety-aware generation. This survey aims to provide a structured conceptual foundation and practical guidance for advancing controllable and reliable video generation models.
CommentsPublished in Transactions on Machine Learning Research (TMLR), 2026. Project page: https://github.com/people-robots/Awesome-Video-Generation-Post-Training
Journal refTransactions on Machine Learning Research, 2026-June, 2026. ISSN 2835-8856