MOJITO:用于统一端到端自动驾驶的模态联合学习
MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving
查看机构详情
- University of Science and Technology of China(中国科学技术大学)
- Li Auto(理想汽车)
- University of New South Wales(新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究针对端到端自动驾驶系统问题,提出基于模态联合学习的MOJITO框架,去除级联接口,执行逐块模态联合注意力更新多模态特征,在数据集上取得新的最优成绩,展现出强大能力。
中文摘要 AI 辅助
端到端自动驾驶系统通常采用级联的两阶段管道,感知阶段将多模态传感器输入压缩为紧凑的上下文,下游规划器据此预测轨迹。我们认为这种单向感知到规划的接口会使传感器输入失去对规划至关重要的细粒度细节,且难以利用现代视觉基础模型的丰富表示。为解决这些问题,我们提出MOJITO,这是一个基于模态联合学习的端到端自动驾驶统一传感器到动作框架。它去除了级联接口,执行逐块模态联合注意力,同时更新动作、图像和激光雷达特征,使规划器在动作生成时能直接访问多模态特征。在NAVSIM v1数据集上达到88.9 PDMS,在更具挑战性的NAVSIM v2数据集上达到88.4 EPDMS,创造了新的最先进水平。广泛实验进一步证明了其强大的可扩展性、指令跟随和多样轨迹生成能力。代码和模型可在指定网址获取。
英文摘要
End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downstream planner predicts trajectories conditioned on this context. We argue that this one-way perception-to-planning interface forces sensor inputs into a compact representation, losing the fine-grained details critical for planning. Moreover, by constraining the planner to this compressed context, it is difficult to leverage the rich representations offered by modern vision foundation models. To address these issues, we propose MOJITO, a unified sensor-to-action framework for end-to-end autonomous driving built on modal joint learning. MOJITO removes the cascaded interface and instead performs block-wise Modal Joint Attention that simultaneously updates action, image, and LiDAR features, allowing the planner to directly access multi-modal features during action generation. MOJITO achieves 88.9 PDMS on the NAVSIM v1 dataset and 88.4 EPDMS on the more challenging NAVSIM v2 dataset, setting a new state-of-the-art. Extensive experiments further demonstrate strong scalability, instruction following, and diverse trajectory generation. Code and models are available at https://github.com/mumucc01/MOJITO.