发表机构
Shanghai Jiao Tong University; Xiamen University(上海交通大学; 厦门大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NeuIDO提出神经算子学习框架,从视觉观测中学习统一内禀动力学表示,实现零样本动力学推断,推动物理信息4D生成迈向世界模型。
AI 中文摘要
世界模型旨在捕捉环境动力学并预测未来轨迹,在具身智能方面展现出日益增长的潜力。物理信息4D生成将物理模拟与3D物体交互预测相结合,为世界模型提供了一条有前景的路径。然而,该范式依赖于人为施加的动力学假设,而非内化世界动力学,因此距离真正的世界模型仍有差距。为弥合这一差距,我们提出NeuIDO,一种新颖的世界动力学建模框架,它从视觉观测中学习统一的内禀动力学表示,推动物理信息4D生成向世界模型迈进。具体而言,我们将世界建模表述为一个神经算子学习问题,并引入两阶段训练策略,以学习从视觉观测分布到内禀动力学分布的可泛化映射。基于这种观测-动力学映射,NeuIDO能够直接从视频中进行零样本动力学推断,并可通过少样本自适应进一步与复杂真实世界动力学对齐。大量实验表明,NeuIDO有效地将多样视觉观测下的内禀动力学统一到共享表示中,并能在新场景中快速推断动力学。
英文摘要
World models aim to capture environmental dynamics and predict future trajectories, showing growing potential for embodied intelligence. Physics-informed 4D generation integrates physical simulation to predict 3D object interactions, offering a promising pathway toward world models. However, this paradigm relies on manually imposed dynamical assumptions rather than internalizing world dynamics, and thus still leaves a gap toward a true world model. To bridge this gap, we propose NeuIDO, a novel world dynamics modeling framework that learns a unified intrinsic dynamics representation from visual observations, advancing physics-informed 4D generation toward a world model. Specifically, we formulate world modeling as a neural operator learning problem and introduce a two-stage training strategy to learn a generalizable mapping from the visual observation distribution to the intrinsic dynamics distribution. Building on this observation-dynamics mapping, NeuIDO enables zero-shot dynamics inference directly from videos and can be further aligned with complex real-world dynamics via few-shot adaptation. Extensive experiments demonstrate that NeuIDO effectively unifies the intrinsic dynamics underlying diverse visual observations into a shared representation and rapidly infers dynamics in novel scenes.
CommentsAccepted by ECCV 2026; Project Page: https://github.com/JiajingLin/NeuIDO