发表机构
UCSD; Adobe(加利福尼亚大学圣迭戈分校; 奥多比公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对主流视频扩散模型未建模像素时间转换的问题,提出LDR方法,在PhyWorld基准上实现更优动态外推,参数更少、速度更快,是首个能泛化至训练分布外的视频世界模型
AI 中文摘要
世界按照其动态规律(即运动定律)演化。然而,主流视频扩散模型大多仅拟合像素,未对像素随时间的转换过程进行建模,因此生成的帧视觉上看似合理,但可能并不准确遵循运动定律。为了仅从像素中捕获动态规律,我们提出了潜动态推理(Latent Dynamics Reasoning,LDR)方法。LDR将潜态转换建模为显式运动学积分,其中低阶动态通过数值积分处理,模型仅回归驱动序列展开的三阶及更高阶残差。为使该积分更好地进行外推,LDR在结构化潜态而非密集卷积特征上运行。我们在PhyWorld(受控白盒物理基准,涵盖匀速运动、抛物线、碰撞、弹跳、逼近5项任务)上验证了LDR,重点关注分布外场景以判断模型是否真正学习到了底层动态。LDR的动态外推效果显著更优:在256²分辨率下,无论是单任务还是联合任务训练,其分布内与分布外误差的差距均比视频扩散基线小20倍以上,同时参数减少26倍,运行速度快143倍。LDR甚至能在严重分布偏移下泛化:例如仅在红色球从左向右运动的数据上训练,它能正确预测蓝色正方形从右向左运动的情况。据我们所知,这是首个能将学习到的动态外推至训练分布之外的视频世界模型。项目页面:this https URL
英文摘要
The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how the pixels transit over time. Thus, they render visually plausible frames but may not accurately obey the laws. To capture the dynamics purely from pixels, we introduce Latent Dynamics Reasoning (LDR). LDR casts the latent transition as an explicit kinematic integration, where the lower-order dynamics are integrated numerically and the model regresses only the third- and higher-order residual that drives the rollout. For this integration to extrapolate better, LDR runs it on a structured latent rather than dense convolutional features. Following PhyWorld, we validate LDR on a controlled white-box physics benchmark spanning five tasks (uniform motion, parabola, collision, bouncing, looming), focusing on out-of-distribution scenarios that reveal whether a model has truly learned the underlying dynamics. LDR extrapolates the learned dynamics far better: the gap between its in- and out-of-distribution error is over 20$\times$ smaller than the video diffusion baseline's, under both single- and joint-task training at 256$^2$ resolution, while using 26$\times$ fewer parameters and running 143$\times$ faster. LDR can even generalize under severe shift: for example, trained only on red balls moving left-to-right, it correctly predicts the motion of a blue square moving right-to-left. To our knowledge, this is the first video world model that extrapolates learned dynamics beyond its training distribution. Project page: https://lat-dyn-reason.github.io/
CommentsProject page: https://lat-dyn-reason.github.io/