arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

S2PD:用于物理和逻辑一致视频生成的串行到并行扩散

S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation

Jeffrey Hu, Daniel Olmeda Reino, Ayush Tewari

arXiv 2610.06847首次发表:更新:

发表机构

University of Cambridge; Toyota Motor Europe(剑桥大学; 丰田汽车欧洲公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

S2PD结合高噪声自回归与低噪声并行扩散,在游戏、物理模拟和真实视频中比双向基线更可靠地遵循规则,并提升时间稳定性与采样效率。

AI 中文摘要

双向视频扩散模型以并行方式对整个视频进行去噪,然而,当在程序化生成器产生的实际上无限分布内数据上训练时,它们仍会违反物理定律和简单的符号规则。我们提出了串行到并行扩散(S2PD),该方法在高噪声水平下执行自回归扩散,然后在低噪声水平下切换到并行扩散。自回归阶段提供了协调相互依赖事件和产生有效状态转换所需的串行计算,而并行阶段则联合细化整个视频,并相比完全串行生成减少了采样时间。我们使用两种架构实现了S2PD:一种从头训练的像素空间扩散变换器,以及一种通过LoRA微调并采用因果注意力的预训练视频模型。在游戏、物理模拟和真实视频中,S2PD比匹配的双向基线更可靠地遵循规则,并且比其他串行方法生成具有更高时间稳定性和采样效率的视频。

英文摘要

Bidirectional video diffusion models denoise entire videos in parallel, yet when trained on effectively unlimited in-distribution data from procedural generators, continue to violate physical laws and simple symbolic rules. We introduce Serial-to-Parallel Diffusion (S2PD), which performs autoregressive diffusion at high noise before switching to parallel diffusion at low noise. The autoregressive phase provides the serial computation needed to coordinate interdependent events and produce valid state transitions while the parallel phase jointly refines the entire video and reduces sampling time relative to fully serial generation. We implement S2PD with two architectures: a pixel-space diffusion transformer trained from scratch and a pretrained video model adapted through LoRA fine-tuning with causal attention. Across games, physical simulations, and real video, S2PD follows rules more reliably than matched bidirectional baselines and generates videos with greater temporal stability and sampling efficiency than other serial methods.

CommentsProject Page: https://jefequien.github.io/S2PD/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑