发表机构
National University of Singapore (NUS); Hunan University; The University of Western Australia (UWA)(新加坡国立大学; 湖南大学; 西澳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对现有VLA自动驾驶模型决策不可靠、多模态交互不足的问题,提出含三类核心组件的多模态交互与多轨迹规划系统,实验显示其在安全推理与场景感知上优于现有系统。
AI 中文摘要
视觉-语言-动作(VLA)模型已成为端到端自动驾驶领域的强大范式,可在统一多模态框架内整合感知、推理与决策。但现有多数VLA模型将端到端自动驾驶形式化为视觉问答任务,导致决策推理不可靠、可解释性差;且无法在异质传感器间建立有效多模态交互,限制了长尾驾驶场景下的鲁棒场景感知与可靠驾驶推理。为此,本文提出一种基于VLA的鲁棒端到端自动驾驶系统,将多模态交互与多轨迹规划优化结合,以实现更可靠、可解释、更安全的驾驶决策。该方法包含三个核心组件:(1)主-辅模态双向交互的亲和引导最优传输;(2)异质模态分布迁移与跨模态交互的分布一致模态迁移;(3)面向长尾驾驶场景的多模态多轨迹规划及感知导向轨迹优化,以生成更优驾驶决策。在开环与闭环数据集上的实验结果表明,与现有驾驶系统相比,本文方法在安全长时程驾驶推理、道路场景感知方面均有提升,凸显了多模态交互与多轨迹规划优化在可扩展VLA系统中的能力。
英文摘要
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for end-to-end autonomous driving by jointly integrating perception, reasoning, and decision making within a unified multimodal framework. However, most existing VLA models formulate end-to-end autonomous driving as a visual question answering task, leading to unreliable and less interpretable decision reasoning. In addition, they fail to establish effective multi-modal interaction across heterogeneous sensors, thereby limiting robust scene perception and reliable driving reasoning in long-tail driving scenarios. To this end, we propose a robust VLA-based end-to-end autonomous driving system that combines multi-modality interaction with multi-trajectory planning and optimization, enabling more reliable, interpretable, and safer driving decisions. Our method comprises three core components: (1) Affinity-Guided Optimal Transport for main-auxiliary modality two-way interaction; (2) Distribution-Consistent Modality Transfer for heterogeneous modality distribution transfer and cross-modal interaction; (3) Multi-modal Multi-Trajectory Planning along with Perception-Oriented Trajectory Refinement for better driving decisions to long-tail driving scenarios. Experimental results in open-loop and closed-loop datasets demonstrate improvements in safety long-horizon driving reasoning and road scene perception over existing driving systems, highlighting the ability of our mutli-modality interaction and multi-trajectory planning and optimization for scalable VLA-based systems.