发表机构
University of Oxford; University of Cambridge; University of Twente; National University of Singapore; University of Bath(牛津大学; 剑桥大学; 特文特大学; 新加坡国立大学; 巴斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Artemis,一种基于几何约束的多智能体驾驶世界模型,通过显式3D状态共享和渐进式记忆更新,解决现有方法缺乏几何约束和静态背景假设的问题,提升视觉保真度和跨视角一致性。
AI 中文摘要
最近的视频世界模型见证了从单智能体到多智能体参与的范式转变,这可以揭示现实世界中更复杂的动态和智能体间的交互。然而,现有方法通常通过交叉注意力采用隐式的智能体间通信,缺乏显式的几何约束和统一的3D状态,从而导致多视角一致性较差,并且在恢复视野外智能体方面存在困难。此外,大多数方法假设静态背景,无法表示不受控制的背景动态。为了解决这些问题,我们提出了Artemis:一种具有显式记忆共享的几何约束多智能体世界模型。从多智能体观测中重建一个显式的3D世界地图,以在智能体之间强制执行统一的3D状态,提供高跨视角一致性。具体而言,开发了一个动作引导的几何注入模块,同时渲染分解的前景-背景控制图,然后通过设计的GeoAdapter块将其注入扩散Transformer。与之前假设仅静态背景的方法相比,我们的GeoAdapter还可以区分不受控制的非智能体动态,这些动态以其自身的多帧历史位置为条件,以提供一致的运动线索。从渐进式视频展开中选择的关键帧用于逐步更新重建的3D世界地图。为了有效捕捉复杂的动态模式,我们策划了一个从CARLA模拟器采样的新数据集,称为MA-CARLA。大量实验证明了我们提出的方法在生成视频的视觉保真度和跨视角一致性方面的优越性。此外,我们的Artemis可以支持同时多模态展开,同时维护2D视频和3D点图,可扩展到超过两个智能体和多摄像头设置。
英文摘要
Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics and cross-agent interaction in the real world. However, existing approaches commonly adopt implicit inter-agent communications via cross attention, which lack explicit geometry constraints and unified 3D state, thereby leading to poor multi-view consistency and struggling with recovering out-of-sight agents. In addition, most of them assume a static background, failing to represent uncontrolled background dynamics. To address these problems, we propose Artemis: a geometry-grounded multi-agent world model with explicit memory sharing. An explicit 3D world map is reconstructed from multi-agent observations to enforce a unified 3D state across agents, offering high cross-view consistency. Specifically, an action-guided geometric injection module is developed to simultaneously render decomposed foreground-background control maps, which are then injected into a diffusion transformer through a designed GeoAdapter block. Compared to previous methods assuming static-only background, our GeoAdapter can also distinguish uncontrolled non-agent dynamics, which are conditioned on their own multi-frame history positions to provide consistent motion cues. Keyframes selected from progressive video rollouts are used to progressively update the reconstructed 3D world maps. To effectively capture complex dynamic patterns, we curate a novel dataset sampled from the CARLA simulator called MA-CARLA. Extensive experiments demonstrate the superiority of our proposed method in terms of visual fidelity and cross-view consistency in the generated videos. In addition, our Artemis can support simultaneous multi-modal rollouts with both 2D video and 3D point map maintenance, scale to scenarios beyond two agents and multi-camera setting.
CommentsThe first three authors contributed equally, and their order was determined by drawing lots. Project Lead: Jiuming Liu. Corresponding Author: Ayush Tewari. Project page: https://liujiuming123.github.io/Artemis/