arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向基于视觉-语言-动作(VLA)的端到端自动驾驶的协同多模态交互

A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving

Jingtao Sun, Xiaohai He, Yike Zhang, Dong Huang, Yaonan Wang, Ajmal Mian, Mike Zheng Shou

arXiv 2608.20890首次发表:更新:

发表机构

National University of Singapore (NUS); Hunan University; The University of Western Australia (UWA)(新加坡国立大学; 湖南大学; 西澳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对现有VLA自动驾驶模型决策不可靠、多模态交互不足的问题,提出含三类核心组件的多模态交互与多轨迹规划系统,实验显示其在安全推理与场景感知上优于现有系统。

AI 中文摘要

视觉-语言-动作(VLA)模型已成为端到端自动驾驶领域的强大范式,可在统一多模态框架内整合感知、推理与决策。但现有多数VLA模型将端到端自动驾驶形式化为视觉问答任务,导致决策推理不可靠、可解释性差;且无法在异质传感器间建立有效多模态交互,限制了长尾驾驶场景下的鲁棒场景感知与可靠驾驶推理。为此,本文提出一种基于VLA的鲁棒端到端自动驾驶系统,将多模态交互与多轨迹规划优化结合,以实现更可靠、可解释、更安全的驾驶决策。该方法包含三个核心组件:(1)主-辅模态双向交互的亲和引导最优传输;(2)异质模态分布迁移与跨模态交互的分布一致模态迁移;(3)面向长尾驾驶场景的多模态多轨迹规划及感知导向轨迹优化,以生成更优驾驶决策。在开环与闭环数据集上的实验结果表明,与现有驾驶系统相比,本文方法在安全长时程驾驶推理、道路场景感知方面均有提升,凸显了多模态交互与多轨迹规划优化在可扩展VLA系统中的能力。

英文摘要

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for end-to-end autonomous driving by jointly integrating perception, reasoning, and decision making within a unified multimodal framework. However, most existing VLA models formulate end-to-end autonomous driving as a visual question answering task, leading to unreliable and less interpretable decision reasoning. In addition, they fail to establish effective multi-modal interaction across heterogeneous sensors, thereby limiting robust scene perception and reliable driving reasoning in long-tail driving scenarios. To this end, we propose a robust VLA-based end-to-end autonomous driving system that combines multi-modality interaction with multi-trajectory planning and optimization, enabling more reliable, interpretable, and safer driving decisions. Our method comprises three core components: (1) Affinity-Guided Optimal Transport for main-auxiliary modality two-way interaction; (2) Distribution-Consistent Modality Transfer for heterogeneous modality distribution transfer and cross-modal interaction; (3) Multi-modal Multi-Trajectory Planning along with Perception-Oriented Trajectory Refinement for better driving decisions to long-tail driving scenarios. Experimental results in open-loop and closed-loop datasets demonstrate improvements in safety long-horizon driving reasoning and road scene perception over existing driving systems, highlighting the ability of our mutli-modality interaction and multi-trajectory planning and optimization for scalable VLA-based systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑