arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05237cs.CVcs.AI

上下文强制:揭示自回归视频扩散中的上下文效应

In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion

Lingxiao Yang, Liu Liu, Moran Li, Han Feng, Wenjian Cao, Jiangning Zhang, Ye Shi

首次发表
浏览论文内容

中文总结 AI 辅助

针对自回归视频扩散模型依赖干净帧导致的时间一致性差等问题,提出In-Context Forcing方法,通过递减噪声水平的上下文实现自适应指导,兼具时间一致性与动态性,还可并行去噪加速推理,在VBench上性能优于SOTA。

中文摘要 AI 辅助

当前少步自回归视频扩散模型依赖所有当前帧去噪步骤的前序完全去噪的干净帧作为上下文,但这些干净帧会泄露过多局部细节,导致模型走捷径,造成时间语义和动态性受损。受扩散即掩码这一视角启发,我们探索噪声上下文对少步自回归生成的影响,然而仅应用相同噪声水平的上下文提供的指导不足,导致时间一致性差。为解决这一难题,我们提出In-Context Forcing(上下文强制),一种利用噪声水平递减的上下文的渐进式自回归范式。通过对远距离帧应用更少掩码、对相邻帧应用更多掩码,该方法提供自适应指导,有效确保稳健的时间一致性和高帧间动态性。此外,通过解耦对前序干净帧的严格依赖,我们的范式支持跨帧并行去噪,在不牺牲性能的情况下实现显著的推理加速。在VBench上开展的大量实验表明,我们的方法在视觉保真度和推理速度上均显著优于当前最优方法。

英文摘要

Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the current frame. However, these clean frames leak excessive local details, which causes the model to take shortcuts, resulting in compromised temporal semantics and dynamics. Inspired by the perspective of diffusion as masking, we explore the impact of noisy contexts on few-step autoregressive generation. Yet, simply applying contexts with the same noise levels provides insufficient guidance, leading to poor temporal consistency. To resolve this dilemma, we introduce In-Context Forcing, a progressive autoregressive paradigm that utilizes contexts with decreasing noise levels. By applying less masking to distant frames and more masking to adjacent ones, this approach provides adaptive guidance, effectively ensuring both robust temporal consistency and high inter-frame dynamics. Furthermore, by decoupling the strict dependence on previous clean frames, our paradigm enables cross-frame parallel denoising, achieving substantial inference acceleration without sacrificing performance. Extensive experiments on VBench demonstrate that our method significantly outperforms state-of-the-art approaches in both visual fidelity and inference speed.

发表机构

  • School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)
  • Tencent Youtu Lab(腾讯云图实验室)
  • College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

↑