物理的隐形之手:当视频扩散模型知道的比它们展示的更多
The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show
浏览论文内容
中文总结 AI 辅助
通过逆向扩散过程探测视频扩散模型的潜在轨迹,发现物理合理性可以从扩散变换器状态中线性解码,准确率达81.27%,表明物理有意义的表示是生成式去噪的副产品。
中文摘要 AI 辅助
现代视频扩散模型生成越来越真实和时间上连贯的视频,这激发了它们作为候选世界模拟器的使用。然而,目前尚不清楚这些模型是否内部编码了物理结构,或者仅仅是复现了训练中看到的运动模式。我们通过沿着对应已知物理合理性的真实视频的潜在轨迹探测视频扩散模型来研究这个问题。为了获得这样的轨迹,我们通过从干净视频潜在变量向后积分学习到的速度场到噪声,近似逆向确定性采样过程,从而访问模型的中间状态和注意力图。利用这些恢复的轨迹,我们表明物理合理性可以从扩散变换器状态中线性解码,在IntPhys和InfLevel上达到约81.27%的平均准确率,并优于专门的表示学习基线如V-JEPA和VideoMAE。令人惊讶的是,这个信号在VAE潜在输入中不存在,而是在去噪变换器内部出现,尽管模型没有使用自监督预测目标进行训练。这些发现表明,物理有意义的表示可以作为生成式去噪的副产品产生。
英文摘要
Modern video diffusion models generate increasingly realistic and temporally coherent videos, motivating their use as candidate world simulators. Yet it remains unclear whether these models internally encode physical structure, or merely reproduce motion patterns seen during training. We study this question by probing video diffusion models along latent trajectories corresponding to real videos with known physical plausibility. To obtain such trajectories, we approximately invert the deterministic sampling process by integrating the learned velocity field backward from a clean video latent to noise, giving access to the model's intermediate states and attention maps. Using these recovered trajectories, we show that physical plausibility is linearly decodable from diffusion transformer states across IntPhys and InfLevel, reaching around 81.27% average accuracy and outperforming dedicated representation-learning baselines such as V-JEPA and VideoMAE. Surprisingly, this signal is absent from the VAE latent input and emerges inside the denoising transformer itself, despite the model not being trained with a self-supervised predictive objective. These findings suggest that physically meaningful representations can arise as a byproduct of generative denoising.
发表机构
- University of Bristol(布里斯托大学)
- McGill University(麦吉尔大学)
- Mila–Quebec AI Institute(魁北克AI研究院)
- Microsoft Research(微软研究院)
- University of Calgary(卡尔加里大学)
机构由 AI 辅助整理,请以论文原文为准。