arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10439cs.CV

流强制:构建用于鲁棒流视频生成的统一训练轨迹

Stream Forcing: Constructing Unified Training Trajectory for Robust Streaming Video Generation

Yueting Zhu, Yuehao Song, Kaicheng Zhang, Bao Tang, Shaoyu Chen, Qian Zhang, Wenyu Liu, Xinggang Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对流视频扩散模型的训练-推理不匹配问题,提出Stream Forcing统一训练框架,通过构建训练轨迹及相关算法,在UCF-101基准测试中FVD分别提升36.6%、27.9%,提升了生成质量与长时序泛化能力。

中文摘要 AI 辅助

流视频生成在世界建模中具有巨大潜力,需在线顺序推断未来帧以形成连续视频流。然而,流视频扩散模型存在根本的训练-推理不匹配问题:推理遵循专门的去噪顺序,而高级训练策略通常需要多样化的噪声水平配置。为解决训练-推理一致性与训练覆盖度之间的权衡,我们将视频扩散采样重新表述为噪声水平上的帧索引随机过程。在该随机过程空间中,我们构建了连续训练轨迹,采样计划沿此轨迹从独立采样逐步演变为与推理一致的采样。我们进一步引入联合校准算法和时间相关采样算法,以确保轨迹平滑性和跨帧相关性。基于这些设计,我们提出了Stream Forcing,一种用于流视频生成的统一训练框架,平衡训练充分性和推理效率。大量实验表明,Stream Forcing在UCF-101基准测试上将FVD指标提升了36.6%,显著提高了生成质量;此外,该方法还能实现对长时序视频生成的鲁棒零样本泛化,在UCF-101基准测试上将FVD指标提升了27.9%。

英文摘要

Streaming video generation holds strong potential for world modeling, where future frames must be inferred online sequentially to form a continuous video stream. However, streaming video diffusion models introduce a fundamental train-inference mismatch: inference follows a specialized denoising order, whereas advanced training strategies typically require diverse noise-level configurations. To address this trade-off between train-inference consistency and training coverage, we reformulate the video diffusion sampling as a frame-indexed stochastic process over noise levels. Within this stochastic process space, we construct a continuous training trajectory along which the sampling schedule progressively evolves from independent sampling to inference-consistent sampling. We further introduce a joint calibration algorithm and a temporal correlative sampling algorithm to ensure trajectory smoothness and cross-frame correlation. Building on these designs, we propose Stream Forcing, a unified training framework for streaming video generation that balances training sufficiency and inference efficiency. Extensive experiments demonstrate that Stream Forcing significantly improves generation quality with a 36.6% FVD improvement on the UCF-101 benchmark. Furthermore, our method facilitates robust zero-shot extrapolation to long-horizon video generation with a 27.9% FVD improvement on the UCF-101 benchmark.

发表机构

  • Huazhong University of Science & Technology(华中科技大学)
  • Horizon Robotics(地平线机器人)
  • Anyverse Dynamics

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑