发表机构
Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory(上海交通大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出免训练的Diff-VF框架,通过三种互补策略及跳跃残差引导,实现高质量长视频生成,在时间一致性与动作多样性上优于现有基线。
AI 中文摘要
近年来,扩散模型在视频生成领域取得了显著进展。然而,现有多数视频扩散模型均以短视频进行训练,当外推至长视频时会出现性能下降,难以在保留多样动作的同时维持长程时间一致性。为生成一致、高质量且动态的长视频,我们提出Diff-VF,这是一种免训练、即插即用且与模型无关的框架,无需修改或微调基础模型即可将现有短视频扩散主干转换为长视频生成器。Diff-VF结合了三种互补策略:混合噪声初始化(HNI)以约束全局语义,加权窗口采样(WWS)以消除窗口间不连续性,以及时间扩展采样(TES)通过随时间步变化的融合建立长程依赖。我们进一步通过跳跃残差引导将Diff-VF扩展至长视频增强,该引导通过随时间步变化的平衡保真度与真实感。VBench-Long评估结果显示,Diff-VF在时间一致性与动作多样性之间取得了比基础模型及近期免训练长视频生成基线(包括FreeNoise、FreeLong和RIFLEx)更优的平衡,同时保持了具有竞争力的逐帧质量。在两个基础模型上的实验证明了其对具有不同时空建模策略的视频扩散模型的适用性,大量 ablation 验证了各组件及超参数的贡献。
英文摘要
Recently, diffusion models have made great progress in video generation. However, most existing video diffusion models are trained with short videos, and degrade when extrapolated to long videos, struggling to maintain long-range temporal coherence while retaining diverse motions. To generate consistent, high-quality and dynamic long videos, we propose Diff-VF, a training-free, plug-and-play and model-agnostic framework that converts existing short-video diffusion backbones into long-video generators without modifying or fine-tuning the base model. Diff-VF couples three complementary strategies: Hybrid Noise Initialization (HNI) to constrain global semantics, Weighted Window Sampling (WWS) to remove inter-window discontinuities, and Temporal Extended Sampling (TES) to establish long-range dependencies with a timestep-varying fusion. We further extend Diff-VF to long-video enhancement via Skip Residual Guidance that balances fidelity and realism through timestep-dependent guidance. VBench-Long evaluation results show that Diff-VF achieves a more favorable balance between temporal coherence and motion diversity than base models and recent training-free long video generation baselines, including FreeNoise, FreeLong, and RIFLEx, while maintaining competitive frame-wise quality. Experiments on two base models demonstrate the applicability to video diffusion models with different spatial-temporal modeling strategies. Extensive ablations validate the contribution of each component and hyperparameters.
CommentsAccepted for publication in ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)