arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MOSH-WM:面向以对象为中心的世界模型的基于掩码的软哈密顿动力学

MOSH-WM: Mask-Grounded Soft-Hamiltonian Dynamics for Object-Centric World Models

Zhekai Wang, Haoxiang Huang, Xiang Liu, Zhikang Chen, Yueqing Sun, Qi Gu, Shiji Zhou, Miao Liu, Sen Cui

arXiv 2608.22750首次发表:更新:

AI 中文总结

该研究提出MOSH-WM模型,通过掩码支撑构建软哈密顿动力学,在OBJ3D、CLEVRER数据集的视频预测任务中,相比基线显著降低误差,且误差积累更慢。

AI 中文摘要

以对象为中心的世界模型通过演化一组实体槽来预测未来视频,但接受动力学监督的变量通常是无约束的视觉特征。我们提出了\method{},一种基于掩码的软哈密顿世界模型,其类位置状态明确依赖于槽所拥有的图像支撑。冻结的视频-槽编码器生成槽和掩码;掩码所拥有支撑的空间矩形成正则状态$Q$,时间差形成$P$,而学习到的能量为有界学习增量提供软方向偏置。与解码器相关的外观和身份单独存储在因果视觉上下文里。门控组合器和有界残差随后将该上下文与传播的相位状态结合,以重构与解码器兼容的槽。在OBJ3D数据集上,给定6个观测帧并对后续30帧进行评估时,\method{}相比最强的以对象为中心的基线,将LPIPS降低了25.0%,空间MSE降低了33.7%。在CLEVRER数据集上,给定6个观测帧并对后续10帧进行评估时,对应降低幅度分别为14.5%和18.7%。按时间范围划分的视觉和对象状态测量结果表明,完整模型在整个30帧闭环滚动过程中误差积累更慢。项目页面:this https URL。

英文摘要

Object-centric world models forecast future videos by evolving a set of entity slots, but the variables receiving dynamics supervision are often unconstrained visual features. We introduce \method{}, a mask-grounded soft-Hamiltonian world model that makes its position-like state explicitly depend on slot-owned image support. A frozen video-slot encoder produces slots and masks; spatial moments of mask-owned support form a canonical state $Q$, temporal differences form $P$, and a learned energy supplies a soft directional bias to a bounded learned increment. Decoder-relevant appearance and identity are stored separately in a causal visual context. A gated composer and bounded residual then combine this context with the propagated phase state to reconstruct decoder-compatible slots. On OBJ3D, given six observed frames and evaluated over the following 30 frames, \method{} reduces LPIPS by 25.0\% and spatial MSE by 33.7\% relative to the strongest object-centric baseline. On CLEVRER, given six observed frames and evaluated over the following ten frames, the corresponding reductions are 14.5\% and 18.7\%. Horizon-resolved visual and object-state measurements show that the complete model accumulates error more slowly throughout the 30-frame closed-loop rollout. Project page:https://github.com/moshwm-anon/-moshwm-anon.github.io.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑