arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27036cs.CVcs.LG

通过视频表示正则化缓解复合误差

Mitigating Compounding Error via Video Representation Regularization

发表机构北京大学
查看机构详情
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

Taiye Chen, Qi Zhang, Yisen Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对视频扩散世界模型自回归生成的复合误差问题,研究发现其与表示维度崩溃相关,提出视频表示正则化方法,在VBench指标上显著优于Diffusion Forcing,提升了长视频生成的鲁棒性。

中文摘要 AI 辅助

基于视频扩散的世界模型可实现用于机器人、自动驾驶和仿真任务的长自回归视频生成,但滑动窗口自回归推理会遭受严重的误差累积,随时间推移降低帧质量。尽管该现象已被广泛观察到,但复合误差的潜在机制以及如何实现稳定的长视界生成仍在很大程度上未得到解决。本文中,我们研究了视频世界模型的内部表示动态,发现复合误差与隐藏表示的维度崩溃紧密耦合。具体而言,当生成漂移开始时,模型表示的有效秩(erank)急剧下降,揭示了表示退化与长期展开不稳定性之间的强关联。此外,我们发现单纯的训练数据扩展无法提升模型对误差漂移的抵抗能力,这与主流的扩展范式相矛盾。为解决该问题,我们提出视频表示正则化,一种用于稳定潜在表示并抑制迭代误差累积的轻量型训练约束。与Diffusion Forcing相比,我们的方法在VBench的美学质量指标上从38.65提升至55.56,在成像质量指标上从44.37提升至72.08。本研究首次建立了自回归视频漂移与模型内部表示之间的联系,采用erank作为误差累积的定量指标,揭示了视频世界模型违反直觉的扩展限制,并提出了一种简单却有效的正则化策略以提升长视频生成的鲁棒性。

英文摘要

Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time. Although this phenomenon has been widely observed, the underlying mechanism of compounding error and how to achieve stable long-horizon generation remain largely unresolved. In this paper, we investigate the internal representation dynamics of video world models and discover that compounding error is tightly coupled with dimensional collapse of hidden representations. Specifically, the effective rank of model representations sharply decreases at the onset of generation drift, revealing a strong connection between representational degradation and long-term rollout instability. Furthermore, we find that pure training data scaling fails to boost model resistance to error drift, contradicting mainstream scaling paradigms. To address this problem, we propose video representation regularization, a lightweight training constraint that stabilizes latent representations and suppresses iterative error accumulation. Compared with Diffusion Forcing, our method achieves improvements from 38.65 to 55.56 and from 44.37 to 72.08 on the Aesthetic Quality and Imaging Quality metrics of VBench. Our work establishes the first connection between autoregressive video drifting and model internal representations, adopts erank as a quantitative metric for error accumulation, reveals counterintuitive scaling limitations for video world models, and presents a simple yet effective regularization strategy to improve long video generation robustness.

↑