视频扩散模型中的序列性差距
The Seriality Gap in Video Diffusion Models
浏览论文内容
中文总结 AI 辅助
研究视频扩散模型在多球动力学实验中的表现,发现其存在序列性差距,即任务所需串行计算与模型去噪循环不匹配,增加有效串行计算的方法可提升性能,还证明去噪步骤在串行推理和模拟任务上有结构障碍。
中文摘要 AI 辅助
当一个球撞击另一个球,然后再撞击另一个球时,视频模型应该预测每次反弹的结果。在多球硬球动力学的控制实验中,我们发现即使提供更多的去噪步骤,标准双向视频扩散的性能也会随着因果链的延长而下降。在没有球-球相互作用的长度匹配单球控制中,这种性能下降基本消失,这表明是依赖事件结构而非视频长度导致了性能下降。在干预研究中,增加有效串行计算的方法能显著提高性能,包括自回归/逐块生成和架构深度。我们将这种模式识别为序列性差距,即要求不断增加串行计算的任务与去噪循环无法提供可扩展串行计算的视频扩散模型之间的不匹配。然后我们证明,对于确定性视频预测,去噪步骤不会在主干之外增加串行计算,这表明在串行推理和模拟任务上视频扩散存在结构障碍。
英文摘要
When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard-sphere dynamics, we find that the performance of standard bidirectional video diffusion degrades as the causal chain lengthens, even when provided more denoising steps. In a length-matched single-ball control, where ball-ball interactions are absent, the degradation largely disappears, isolating dependent-event structure rather than video length as the cause. Across intervention studies, methods that increase effective serial computation improve performance disproportionately, including autoregressive/blockwise generation and architectural depth. We identify this pattern as the seriality gap: a mismatch between tasks requiring growing serial computation and video diffusion models whose denoising loop does not provide scalable serial compute. We then prove that, for deterministic video prediction, denoising steps do not add serial computation beyond the backbone, indicating a structural obstacle for video diffusion on serial reasoning and simulation tasks.
发表机构
- UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。