GaussianDWM++:用于统一场景理解、编辑与多模态生成的语言接地3D高斯驾驶世界模型
GaussianDWM++: Language-Grounded 3D Gaussian Driving World Model for Unified Scene Understanding, Editing, and Multi-Modal Generation
浏览论文内容
中文总结 AI 辅助
该研究提出GaussianDWM++,通过基础特征高斯分词器等模块构建统一框架,在多类驾驶任务中实现场景理解、编辑与多模态生成,性能达当前最优。
中文摘要 AI 辅助
驾驶世界模型(Driving World Models,DWMs)近期随生成模型快速发展,但多数现有方法主要聚焦条件场景生成,缺乏显式3D场景理解、语言接地推理及可控4D编辑能力。此外,常用的点云、占据或鸟瞰图(BEV)表示难以实现文本信息与底层3D场景结构的细粒度对齐。为解决这些局限,我们提出一种基础特征高斯驾驶世界模型,在单一框架内统一场景理解、语言接地推理、可控4D编辑与多模态生成。具体而言,我们引入基础特征高斯分词器,直接将Qwen/SigLIP视觉-语言特征提炼为3D高斯基元,构建紧凑的开放词汇高斯语义场;进一步设计几何感知高斯适配器,结合重要性感知分层选择与文本条件感知器(Perceiver)风格交叉注意力,将密集高斯基元聚合为紧凑世界token;为提升表示兼容性,引入基于KL的高斯-图像分布对齐目标,使高斯世界token与基础图像token对齐。基于对齐后的高斯表示,我们的框架还支持指令可控的场景编辑,包括天气条件生成与动态车辆操控。在更广泛的驾驶基准上开展的大量实验表明,我们的方法在场景理解、视觉接地、规划导向推理及可控4D生成任务中均达到了当前最优性能。我们将在Github上公开发布代码与数据集。
英文摘要
Driving World Models (DWMs) have recently advanced rapidly with generative models, yet most existing methods mainly focus on conditional scene generation and lack explicit 3D scene understanding, language-grounded reasoning, and controllable 4D editing capabilities. Moreover, commonly used point cloud, occupancy, or BEV representations make it difficult to achieve fine-grained alignment between textual information and the underlying 3D scene structure. To address these limitations, we propose a foundation-feature Gaussian driving world model that unifies scene understanding, language-grounded reasoning, controllable 4D editing, and multi-modal generation within a single framework. Specifically, we introduce a foundation-feature Gaussian tokenizer that directly distills Qwen/SigLIP visual-language features into 3D Gaussian primitives, building a compact open-vocabulary Gaussian semantic field. We further design a geometry-aware Gaussian adapter that combines importance-aware hierarchical selection with text-conditioned Perceiver-style cross-attention to aggregate dense Gaussian primitives into compact world tokens. To improve representation compatibility, we introduce a KL-based Gaussian--image distribution alignment objective that aligns Gaussian world tokens with foundation image tokens. Based on the aligned Gaussian representation, our framework further supports instruction-controllable scene editing, including weather-conditioned generation and dynamic vehicle manipulation. Extensive experiments on broader driving benchmarks demonstrate that our method achieves state-of-the-art performance across scene understanding, visual grounding, planning-oriented reasoning, and controllable 4D generation tasks. We will release the code and datasets publicly on Github.
发表机构
- School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院)
- Tsinghua University(清华大学)
- Chongqing Afari Intelligent Drive Co., Ltd.(重庆阿法瑞智能驾驶有限公司)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。