arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

4DR360:用于4D雷达-相机全场景感知中联合3D检测和占用预测的状态推理

4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception

Xiaokai Bai, Lianqing Zheng, Runwei Guan, Songkai Wang, Siyuan Cao, Hui-liang Shen

arXiv 2607.09629首次发表:更新:

AI 中文总结

针对4D雷达-相机全场景感知,提出\method框架,遵循跨模态状态推理范式,通过状态引导的BEV增强和多普勒引导的时间融合进行联合3D检测和占用预测,扩展数据集并实验,提升多任务学习效果。

AI 中文摘要

可靠的自动驾驶需要将前景物体与密集语义布局相结合的全场景感知。4D毫米波雷达虽已成为强大且经济的传感器,但其稀疏回波使雷达-相机融合对全面场景理解必不可少。现有方法主要优化检测,双任务系统交互有限。为此提出\method框架用于360°全场景感知,将语义占用建模为持久场景状态。该框架遵循跨模态状态推理范式,通过阶段进行粗到细的特征聚合来建模和传播占用状态。具体包括状态引导的BEV增强和多普勒引导的时间融合。还扩展数据集并在统一协议下实验,涵盖精度、鲁棒性等方面,代码和标签接受后发布。

英文摘要

Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millimeter-wave radar has emerged as a robust and affordable sensor, yet its sparse returns make radar-camera fusion necessary for comprehensive scene understanding. Existing radar-camera methods mainly optimize detection, while dual-task systems usually decode boxes and occupancy with limited interaction. To address this gap and advance radar-based multi-task learning, we propose \method, a 4D radar-camera framework for 360$^\circ$ full-scene perception, which models semantic occupancy as a persistent scene state rather than a terminal output. \method{} follows a cross-modal state reasoning paradigm, where the occupancy state is modeled and propagated through stages for coarse-to-fine feature aggregation. Specifically, State-guided BEV Enhancement (SBE) strengthens intra-frame BEV representation, while Doppler-guided Temporal Fusion (DTF) preserves state evidence over longer temporal horizons. Beyond the model, we further extend ManTruckScenes with satellite-map-based generated occupancy labels and pair it with OmniHD-Scenes in a unified cross-dataset detection-and-occupancy protocol. The resulting experiments cover accuracy, robustness, ablation, and efficiency under one radar-camera multi-task evaluation framework. Code and labels will be released upon acceptance.

Comments5 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑