发表机构
The Chinese University of Hong Kong, Shenzhen; Tencent Hunyuan(香港中文大学(深圳); 腾讯混元)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出条件残差预测(CRP)方案,无需双向教师模型,训练出20亿参数的因果视频模型Optica,大幅缩小与双向模型的质量差距,仅用约1500万训练视频便生成高质量5秒480p视频。
AI 中文摘要
因果视频扩散模型以自回归方式生成视频,适用于流式、交互式及长视频生成。然而在标准训练下,其生成质量常低于同等规模的双向模型。现有诸多方法通过从预训练双向教师模型初始化或蒸馏该模型来缩小此差距。本文则从图像模型初始化训练因果模型,全程不使用双向视频模型,该路径既无需大型双向教师模型,也无需复杂的蒸馏流程,更简单且可扩展性更强。研究发现,基于真实历史训练的因果模型会严重依赖该历史,导致推理时自身生成历史的错误会向前传播。本文假设这种依赖大多不必要,因为当前输入已能确定历史提供的大部分信息,进而提出条件残差预测(Conditional Residual Prediction, CRP)这一降低模型对条件依赖的简单方案:模型先在无条件的情况下预测目标,条件仅在此预测基础上添加残差。将CRP应用于历史时,模型会尽可能从当前预测每个视频块,仅用历史补充当前无法提供的内容。受控实验显示,CRP几乎将同等设置下与双向模型的6.14分差距缩小。扩展该方案后,本文训练出Optica——一个20亿参数的因果视频模型,可自回归生成5秒480p视频,仅用约1500万训练视频便在VBench上达到82.78分。
英文摘要
Causal video diffusion models generate video autoregressively, which suits streaming, interactive, and long-video generation. Under standard training, however, they often yield lower generation quality than bidirectional models of the same size. Many existing approaches address this gap by initializing from or distilling a pretrained bidirectional teacher. We instead train a causal model from an image-model initialization, with no bidirectional video model at any stage. Because this path requires neither a large bidirectional teacher nor a complex distillation pipeline, it is simpler and more scalable. On this path, we find that a causal model trained on ground-truth history becomes strongly dependent on it, so that at inference errors in its own generated history propagate forward. We hypothesize that much of this dependence is unnecessary, because the current input already determines much of what the history provides. We propose Conditional Residual Prediction (CRP), a simple recipe for reducing a model's reliance on a condition: the model first predicts the target without the condition, and the condition may only add a residual on top of this prediction. Applied to history, CRP makes the model predict each chunk from the present as far as it can and use the past only for what the present cannot supply. In controlled experiments, CRP nearly closes the 6.14-point gap to a bidirectional model trained under the same setup. Scaling this recipe, we train Optica, a 2B-parameter causal video model that autoregressively generates 5-second 480p videos and reaches 82.78 on VBench with only about 15M training videos.
Comments26 pages, 6 figures, 4 tables