arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MATS:一种用于自动驾驶中3D感知的新型多模态多任务学习框架

MATS: A novel multi-modality multi-task learning framework for 3D perception in autonomous driving

Junchen Huo, Wanming Hao, Song Wang, Enqing Chen, Shouyi Yang, Guanghui Wang

arXiv 2607.24224首次发表:更新:

发表机构

School of Electrical and Information Engineering, Zhengzhou University; Tianping College of Suzhou University of Science and Technology; Toronto Metropolitan University(郑州大学电气与信息工程学院; 苏州科技大学天平学院; 多伦多都会大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自动驾驶3D感知,提出MATS多模态多任务学习框架,通过模态自适应BEV融合和特定任务MoE模块,在nuScenes基准测试中,多模态输入下显著优于现有技术,单任务也优于基线。

AI 中文摘要

来自不同传感器的多模态数据为3D感知提供了丰富的互补信息,是可靠自动驾驶系统的重要组成部分。当前研究通常设计复杂的融合策略,在统一的鸟瞰图(BEV)特征图上整合多模态数据信息以联合学习多个感知任务,但单一特征图难以满足各任务需求,导致感知性能受限。本文提出MATS,一种具有模态自适应BEV融合和特定任务专家混合(MoE)的新型多模态多任务学习方法用于3D感知。设计了简单的模态自适应BEV融合模块,通过建模全局跨模态依赖自适应重新校准BEV特征,为不同感知任务生成多样的BEV特征图。还提出特定任务的MoE模块解耦任务,使网络能为每个特定任务自动选择合适的BEV特征候选。在大规模基准nuScenes上进行大量实验,结果表明该方法在多模态输入数据下显著优于现有技术,在单任务上也明显优于基线。代码和训练模型将在发表后提供。

英文摘要

Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex fusion strategies to integrate information from multimodal data on a unified bird's-eye-view (BEV) feature map for the joint learning of multiple perception tasks. However, such a single feature map hardly carries sufficient information to simultaneously meet the requirements of various perception tasks, leading to a very limited perception performance. To mitigate this limitation, this paper proposes MATS, a novel multi-modality multi-task learning approach with modality-adaptive BEV fusion and task-specific Mixture-of-Experts (MoE) for 3D perception. Specifically, a simple modality-adaptive BEV fusion module is designed to adaptively recalibrate the BEV features by modeling the global cross-modality dependencies, generating diverse BEV feature maps for various perception tasks. For joint multi-task learning, this paper proposes a task-specific MoE module to decouple the tasks and enable the network to automatically choose the appropriate BEV feature candidates for each specific task. To validate the effectiveness of the proposed approach, we conduct extensive experiments on the large-scale benchmark nuScenes. With the camera- and LiDAR-modality input data, the proposed approach outperforms the state-of-the-art (SOTA) by a significant margin. Furthermore, the experimental results on the single tasks show that the proposed approach significantly outperforms the baselines. The code and trained models will be available upon publication.

Comments12 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑