发表机构
Huawei Noah’s Ark Lab(华为诺亚方舟实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GeoWM提出一种直接预测未来场景几何的几何世界模型,无需递归展开,利用几何基础模型和流匹配变换器,在四个数据集上优于现有模型并大幅降低推理时间。
AI 中文摘要
对3D场景几何及其随时间演变的建模对于自动驾驶和机器人技术至关重要。一种常见的范式是使用世界模型来预测环境的未来图像或潜在表示,随后从这些预测中恢复几何信息。然而,这种范式并未显式地建模几何结构,且通常依赖递归展开以达到更长的预测范围,导致误差累积和计算成本增加。为解决这些局限性,我们提出了GeoWM,一种几何世界模型,它直接在指定的未来时间范围预测未来的场景几何,无需递归展开。关键思想是利用几何基础模型将观测到的RGB帧转换为几何历史,该历史条件化一个流匹配变换器,以预测指定未来时间范围的场景几何。我们进一步表明,一个轻量级的相机运动预测器能够准确估计未来视点,并且将观测到的几何投影到预测视点中,为未来几何预测提供了有效的几何先验。在涵盖城市驾驶、空中飞行和动态操作四个数据集上的大量实验表明,GeoWM在预测深度、相机姿态和3D场景几何方面优于所评估的世界模型,同时在更长的时间范围上大幅减少了推理时间。
英文摘要
Modeling 3D scene geometry and its evolution over time is essential for autonomous driving and robotics. A common paradigm is to use world models to predict future images or latent representations of the environment and subsequently recover geometry from these predictions. However, this paradigm does not explicitly model geometric structure and typically relies on recursive rollouts to reach longer prediction horizons, leading to error accumulation and increasing computational cost. To address these limitations, we present GeoWM, a geometry world model that directly forecasts future scene geometry at specified future horizons without recursive rollout. The key idea is to leverage a geometry foundation model to transform observed RGB frames into a geometric history, which conditions a flow-matching transformer to predict the scene geometry at a specified future horizon. We further show that a lightweight camera-motion predictor can accurately estimate the future viewpoint, and that projecting the observed geometry into the predicted viewpoint provides an effective geometric prior for future geometry forecasting. Extensive experiments on four datasets spanning urban driving, aerial flight, and dynamic manipulation demonstrate that GeoWM outperforms the evaluated world models in forecasting depth, camera pose, and 3D scene geometry, while substantially reducing inference time at longer horizons.