arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将物理先验知识蒸馏为流式世界模型

Distilling Physical Priors into Streaming World Models

Liangliang Zhao, Junying Wang, Danni Yang, Yifan Chang, Bin Fu, Yu Qiao, Bowen Zhou, Yihao Liu

arXiv 2608.07981首次发表:更新:

AI 中文总结

本文提出PhyS三阶段框架,构建120K物理交互数据集,经微调、蒸馏及在线强化学习结合TCR方法,提升流式世界模型的物理一致性,在PhysicsIQ等基准上取得显著性能提升。

AI 中文摘要

流式世界模型可在线预测未来视觉状态,同时在长时序范围内保持物理一致性的动力学特性。然而,这类模型的生成序列常违反基本物理约束。现有常见方法是将预训练的双向DiT蒸馏为少步因果生成器,但该范式存在两个核心局限:通用双向教师模型从面向视觉的预训练中获取的物理先验有限,且在双向到因果的蒸馏过程中,有限的先验会进一步流失。本文提出PhyS,这是一个将物理先验知识蒸馏为流式世界模型的三阶段框架。为从真实世界交互中获取物理先验,我们构建了PhyS-120K数据集,包含120K个真实世界物理交互视频,覆盖刚体动力学、软体变形、流体现象及相变,每个视频均标注了物体属性和因果状态转换的结构化描述。通过感知物理的监督微调,将物理先验注入到140亿参数的双向DiT教师模型中,随后将其蒸馏为13亿参数的轻量型因果DiT,用于少步自回归流式生成。最后,我们采用在线强化学习激励蒸馏模型生成物理合理的序列,并进一步提出时间信用路由(TCR)以解决时间信用分配问题:TCR在重叠时间窗口内评估物理一致性,并将得到的组相对优势分配给时间对齐的去噪动作。在PhysicsIQ基准上,PhyS使Wan2.1-14B教师模型的性能提升18.2%,使Self Forcing、Rolling Forcing和Causal Forcing方法分别提升23.7%、14.8%和31.4%,同时在感知物理的视频基准VideoPhy、VideoPhy2和PhyGenBench上也取得了性能提升。相关数据集、代码及更多样例视频可在项目页面获取。

英文摘要

Streaming world models predict future visual states online while maintaining physically coherent dynamics over long horizons. However, their rollouts often violate basic physical constraints. A common approach distills pretrained bidirectional DiTs into few-step causal generators. However, this paradigm suffers from two fundamental limitations: generic bidirectional teachers acquire limited physical priors from visually oriented pretraining, and the limited priors suffer further loss during bidirectional-to-causal distillation. We present PhyS, a three-stage framework for distilling physical priors into streaming world models. To acquire physical priors from real-world interactions, we construct PhyS-120K, a dataset of 120K real-world physical-interaction videos spanning rigid-body dynamics, soft-body deformation, fluid phenomena, and phase transitions. Each video is annotated with structured descriptions of object properties and causal state transitions. Physics-aware supervised fine-tuning injects the physical priors into a bidirectional 14B DiT teacher, which we then distill into a lightweight 1.3B causal DiT for few-step autoregressive streaming generation. Finally, we use online reinforcement learning to incentivize the distilled model to generate physically plausible rollouts and further propose Temporal Credit Routing (TCR) to address temporal credit assignment. TCR evaluates physical consistency over overlapping temporal windows and routes the resulting group-relative advantages to temporally aligned denoising actions. On PhysicsIQ, PhyS improves the Wan2.1-14B teacher by 18.2\% and the Self Forcing, Rolling Forcing, and Causal Forcing by 23.7\%, 14.8\%, and 31.4\%, respectively. Results also improve the physics-aware video benchmarks VideoPhy, VideoPhy2, and PhyGenBench. The dataset, code, and more sample videos are available on our Project Page.

Comments9 pages, 7 figures. Project page: https://lyongo.github.io/PhyS/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑