arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考流式视频扩散模型:上下文、执行与训练

Rethinking Streaming Video Diffusion Model: Context, Execution, and Training

Hongchen Zhang

arXiv 2609.22283首次发表:更新:

发表机构

University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出统一分析框架研究流式视频扩散设计空间,发现渐进历史策略在VBench上得分最高且加速1.57-2.83倍,并证明LoRA适配可大幅减少参数,表明完全去噪历史并非必要。

AI 中文摘要

理解流式视频扩散的设计空间对于探索其在生成质量和计算效率方面的潜力至关重要。我们开发了一个统一的分析框架,该框架将模型和采样器选择、历史条件、执行调度以及训练策略联系起来。该框架适应了广泛的因果上下文选择策略家族,并使其计算依赖关系和训练-推理对齐明确化。在此设计空间内,我们研究了三种代表性策略:干净历史、同级历史和渐进历史。在完整的VBench提示集上,同级历史和渐进历史分别取得了85.24和85.60的聚合分数,而干净历史参考为84.45。长视频比较进一步显示,使用渐进历史时,主体一致性和运动连贯性有所改善。通过允许多个去噪节点同时处理,渐进历史流水线在我们评估的条件下实现了1.57至2.83倍的稳态DiT加速。我们还发现,对DMD伪评分网络进行LoRA适配,仅使用全参数适配所需可训练伪评分参数的2.15%,即可提高生成质量。综上所述,这些发现表明,完全去噪的历史并非高质量流式生成的先决条件,并激励了历史条件、执行和训练的联合设计。

英文摘要

Understanding the design space of streaming video diffusion is essential to exploring its potential for generation quality and computational efficiency. We develop a unified analytical framework that relates model and sampler choices, historical conditioning, execution scheduling, and training strategies. The framework accommodates a broad family of causal context-selection policies and makes their computational dependencies and training-inference alignment explicit. Within this design space, we study three representative policies: clean, same-level, and progressive history. On the full VBench prompt set, same-level and progressive history achieve aggregate scores of 85.24 and 85.60, respectively, compared with 84.45 for the clean-history reference. Long-video comparisons further show improved subject consistency and more coherent motion with progressive history. By allowing multiple denoising nodes to be processed together, progressive-history pipelining achieves $1.57$-$2.83\times$ steady-state DiT speedups under our evaluated conditions. We additionally find that LoRA adaptation of the DMD fake-score network improves generation quality using only 2.15% as many trainable fake-score parameters as full-parameter adaptation. Together, these findings show that fully denoised history is not a prerequisite for high-quality streaming generation and motivate the joint design of historical conditioning, execution, and training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑