发表机构
Brown University; Massachusetts Institute of Technology; Meta(布朗大学; 麻省理工学院; Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对流视频生成中传统方法的不足,提出Ms. Forcing范式,通过多尺度分块和注意力调整空间粒度,引入均匀噪声水平DMD减少不匹配,该方法提升了效率,在单个H200 GPU上性能显著优于Rolling Forcing。
AI 中文摘要
流视频扩散模型在交互式和动态世界模拟方面取得了显著进展,但传统的下一帧生成的嵌套自回归和去噪循环阻碍了实时部署。近期的滚动窗口方法在不同噪声水平的多个连续帧上进行去噪流水线操作,提高了吞吐量和长时稳定性。然而,它们以相同的精细空间粒度对每个状态进行令牌化,在联合去噪窗口中留下了大量与噪声相关的冗余。我们提出了Ms. Forcing,一种高效的流视频生成范式,它根据每个状态的噪声水平调整空间粒度。其多尺度分块(MSP)为噪声更大的状态分配更粗的块,将活动窗口令牌数减少45%,而多尺度自注意力(MSSA)使可见非汇键和值的密度与每个查询尺度相匹配,以进一步降低注意力成本。由于两个调度都由窗口位置固定,Ms. Forcing保留了一个静态的、硬件友好的计算图。我们还引入了均匀噪声水平DMD(H-DMD),它从共享相同源噪声水平的干净预测中组装每个虚假视频,从而减少DMD训练序列与推理时展开之间的不匹配。多尺度设计有助于抵消通过重叠窗口进行反向传播的额外训练成本。我们进行了定量和定性实验,结果表明Ms. Forcing在单个H200 GPU上达到22.84 FPS,比Rolling Forcing快39.6%,同时在短视频和长视频生成设置中显著提高了VBench分数。
英文摘要
Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoising loops of conventional next-frame generation hinder real-time deployment. Recent rolling-window methods pipeline denoising across multiple consecutive frames at different noise levels, improving throughput and long-horizon stability. However, they tokenize every state at the same fine spatial granularity, leaving substantial noise-dependent redundancy in the joint denoising window. We propose Ms.Forcing, an efficient streaming video generation paradigm that adapts spatial granularity to each state's noise level. Its Multi-Scale Patchification (MSP) assigns coarser patches to noisier states, reducing the active-window token count by 45%, while Multi-Scale Self-Attention (MSSA) matches the density of visible non-sink keys and values to each query scale to further reduce attention cost. Because both schedules are fixed by window position, Ms.Forcing retains a static, hardware-friendly computation graph. We further introduce Homogeneous-Noise-Level DMD (H-DMD), which assembles each fake video from clean predictions sharing the same source noise level, thereby reducing the mismatch between DMD training sequences and inference-time rollouts. The multi-scale design helps offset the additional training cost of backpropagating through overlapping windows. We include both quantitative and qualitative experiments to show that Ms.Forcing reaches 22.84 FPS on a single H200 GPU, 39.6% faster than Rolling Forcing, while significantly improving VBench scores in both short video and long video generation setting.