4DGS-WAM:基于4D高斯溅射的以对象为中心的世界动作模型,连接过去与未来
4DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting
- The Chinese University of Hong Kong(香港中文大学)
- Shanghai Academy of AI for Science(上海人工智能科学研究院)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出基于4D高斯溅射的以对象为中心的世界动作模型4DGS-WAM,将2D观测转为持久4D表示,复用静态内容以聚焦动态对象演化,在KITTI-MOT上完成短程预测与过去重建评估。
AI中文摘要:
当前的世界动作模型(WAM)通常基于2D视觉数据运行,这些模型可实现出色的视觉质量,但缺乏针对单个对象的显式空间结构,且会重复处理冗余的背景内容。尽管点云可在3D空间中表示世界,但它们在不同视点间的对齐和累积可能存在困难。在本文中,我们利用显式的4D高斯溅射(4DGS)表示,该表示可分别对场景中的动态对象和静态背景进行建模。对于动态对象,我们使用策略模型预测未来的智能体动作,并使用世界模型预测其观测到的高斯溅射的变换;静态背景无需为未来状态重新生成,因为其中大部分已在过去的帧中被观测到。这形成了一个以对象为中心的世界动作模型,我们将其命名为4DGS-WAM。该模型将2D观测提升为持久的4D表示,以便在未来预测过程中复用先前观测到的静态内容,从而使未来状态外推可以专注于建模动态对象的演化。在KITTI-MOT数据集上进行的实验评估了短程预测和过去重建任务。
英文摘要:
Current world action models (WAMs) typically operate on 2D visual data. These models can achieve exceptional visual quality, but they lack explicit spatial structure for individual objects and repeatedly process redundant background content. Although point clouds can represent the world in 3D space, they can be difficult to align and accumulate across viewpoints. In this paper, we leverage an explicit 4D Gaussian Splatting (4DGS) representation that separately models dynamic objects and the static background of a scene. For dynamic objects, we use a policy model to predict future actor actions and a world model to predict transformations of their observed Gaussian splats. The static background need not be regenerated for future states, as much of it has already been observed in past frames. This forms an object-centric world action model, which we name 4DGS-WAM. It lifts 2D observations into a persistent 4D representation so that previously observed static content can be reused during future prediction. Future-state extrapolation can then focus on modeling the evolution of dynamic objects. Experiments on KITTI-MOT evaluate short-horizon prediction and past reconstruction.