arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自动驾驶中鸟瞰视角分割的变分推理

Variational Inference for Bird's Eye View Segmentation in Autonomous Driving

Jingyue Shi, Huaicheng Li, Junhui Zhao, Yanxiang Jiang

arXiv 2607.14710首次发表:更新:

发表机构

School of Electronic and Information Engineering, Beijing Jiaotong University; School of Information Science and Engineering, Southeast University(北京交通大学电子信息工程学院; 东南大学信息科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自动驾驶中鸟瞰视角分割难题,提出基于变压器的变分流变换网络TVB,通过后验BEV监督学习映射,结合条件变分自编码器、归一化流及注意力融合模块,在多摄像头视图BEV分割等方面性能优越。

AI 中文摘要

鸟瞰视角(BEV)已成为自动驾驶环境感知的关键方法,为车辆提供统一空间表示。然而,有效融合多摄像头传感器数据并在复杂外部驾驶环境中运行仍是挑战。为缓解此问题,我们在变分推理框架中重塑BEV分割问题。提出基于变压器的变分流变换网络TVB,训练时通过后验BEV监督隐式学习从多摄像头视图到统一规范BEV图的映射。TVB以条件变分自编码器为骨干,生成多个BEV图候选。通过整合归一化流增强生成图的真实感,设计BEV注意力融合模块自适应整合候选图。在nuScenes和OPV2V数据集上的实验表明,该方法在多摄像头视图BEV分割和车道环境感知中性能优越。

英文摘要

The bird's eye view (BEV) has emerged as a pivotal approach for environmental perception in autonomous driving, providing a unified spatial representation for vehicles. Nevertheless, despite BEV's significance in addressing the challenges inherent to autonomous driving, effectively fusing data from multiple camera sensors and operating in complex external driving environments remains a considerable challenge. To mitigate this issue, we recast the BEV segmentation problem within a variational inference framework. In this paper, we propose a novel transformer-based variational flow transformation network for BEV segmentation, denoted as TVB. Our architecture implicitly learns the mapping from multiple camera views to a unified canonical BEV map during training by exploiting posterior BEV supervision. TVB employs a conditional variational auto encoder (CVAE) as its backbone and produces multiple BEV map candidates. To augment the realism of the generated BEV maps, we integrate normalizing flows into the map generation process, enabling the construction of more complex and expressive probability distributions. Furthermore, we design a BEV-attention fusion (BAF) module that harnesses attention mechanisms to adaptively integrate the multiple candidate BEV maps. Experimental results, evaluated on both the nuScenes and OPV2Vdatasets, demonstrate that our proposed method achieves superior performance in multi-camera view BEV segmentation and lane environment perception.

Comments13 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑