发表机构
Texas Tech University; University of Michigan(德克萨斯理工大学; 密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MM-BEV是一种遵循「计算关键位置与时刻」的实时多模态BEV系统,通过四种机制优化,在nuScenes数据集和Clearpath Husky A300平台上显著降低延迟,且性能损失极小。
AI 中文摘要
多模态鸟瞰图(BEV)感知融合了激光雷达(LiDAR)的深度精度与相机的密集语义,但高昂的计算成本和不完善的感知条件使其难以实时部署。现有方法大多压缩单个检测器,却忽略了三个优化机会:相机与激光雷达输入中的结构化稀疏性、模态间的时序失配,以及许多检测到的物体不会影响规划器的即时行动。本文提出MM-BEV,一种遵循「计算关键的位置与时刻」这一简单原则的实时多模态BEV系统。MM-BEV将感知工作分为两类:针对自车制动距离内、碰撞时间(TTC)短的安全关键物体的强制工作,以及针对紧迫性较低区域的可选工作。该系统在计算资源紧张时优先处理强制工作,减少或舍弃可选工作。MM-BEV整合了四种机制:(1)基于前一帧运动外推检测结果的关键度排序时序感兴趣区域(ROI)选择器;(2)采用上下文自适应分辨率的共享形状相机裁剪与ROI感知激光雷达体素化的稀疏ROI感知特征提取;(3)根据场景动态与TTC调整激光雷达扫描、图像分辨率和关键帧的延迟感知协调器;(4)将感知与推理解耦并跳过过时帧的异步调度器。在nuScenes数据集上,MM-BEV将推理延迟降低1.96倍,端到端延迟降低2.93倍,几何关键召回率无损失,安全关键召回率仅下降0.2个百分点。在搭载Ouster-128激光雷达、BEV相机和Jetson AGX Orin的Clearpath Husky A300平台上,MM-BEV进一步将平均延迟降低2.11倍,展现出在现实自主系统中的应用潜力。
英文摘要
Multimodal bird's-eye-view (BEV) perception combines LiDAR depth accuracy with dense camera semantics, but its high computational cost and imperfect sensing conditions make real-time deployment challenging. Existing methods largely compress individual detectors and overlook three opportunities: structured sparsity within camera and LiDAR inputs, timing misalignment between modalities, and the fact that many detected objects do not affect the planner's immediate action. We present MM-BEV, a real-time multimodal BEV system guided by a simple principle: compute where and when it matters. MM-BEV divides perception into mandatory work for safety-critical objects within braking distance of the ego vehicle and with short time-to-collision (TTC), and optional work for less urgent regions. It prioritizes mandatory work and reduces or sheds optional work under tight compute budgets. MM-BEV integrates four mechanisms: (1) a criticality-ranked temporal ROI selector based on motion-extrapolated detections from prior frames; (2) sparse, ROI-aware feature extraction using shared-shape camera crops at context-adaptive resolution and ROI-aware LiDAR voxelization; (3) a latency-aware coordinator that adapts LiDAR sweeps, image resolution, and keyframes according to scene dynamics and TTC; and (4) an asynchronous scheduler that decouples sensing from inference and skips stale frames. On nuScenes, MM-BEV reduces inference latency by 1.96x and end-to-end latency by 2.93x, with no loss in geometry-critical recall and only a 0.2 percentage-point drop in safety-critical recall. On a Clearpath Husky A300 equipped with an Ouster-128 LiDAR, BEV cameras, and a Jetson AGX Orin, MM-BEV further reduces mean latency by 2.11x, demonstrating its potential for real-world autonomous systems.
Comments12 pages, 20 figures