发表机构
University of South Florida; University of Central Florida(南佛罗里达大学; 中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AquaBEV是基于3D声呐监督的单目水下BEV占用模型,经受控基准测试,其可见IoU、观测IoU较最强迁移基线分别提升4.0%、4.3%。
AI 中文摘要
自主水下机器人广泛应用于勘探、监测与检查,其安全导航依赖于对周围空闲与占用空间的理解,鸟瞰图(BEV)占用正是这样的表示,但仅从单张水下RGB图像预测该表示十分困难,因为仅从外观获取的几何线索有限且不可靠。3D成像声呐提供互补的几何测量以监督该任务。我们提出AquaBEV,这是一种单目水下占用模型,在训练过程中使用配对的3D成像声呐作为几何监督,从单张RGB图像预测局部BEV占用。AquaBEV将视觉特征映射到无需校准的极坐标表示,并沿距离维度应用因果解码,之后在笛卡尔BEV坐标中重建预测结果。我们建立了一个受控水下占用基准,在统一协议下将代表性占用方法适配到相同的RGB到声呐任务中。AquaBEV取得了31.4的可见IoU和38.6的观测IoU,比最强的迁移基线分别实现了4.0%和4.3%的相对提升。
英文摘要
Autonomous underwater robots are widely used for exploration, monitoring, and inspection, where safe navigation depends on understanding the surrounding free and occupied space. Bird's eye view (BEV) occupancy provides such a representation, but predicting it from a single underwater RGB image is difficult due to limited, unreliable geometric cues from appearance alone. 3D imaging sonar offers complementary geometric measurements to supervise this task. We introduce AquaBEV, a monocular underwater occupancy model that predicts local BEV occupancy from a single RGB image, using paired 3D imaging sonar as geometric supervision during training. AquaBEV maps visual features into a calibration free polar representation and applies causal decoding along the range dimension before reconstructing the prediction in Cartesian BEV coordinates. A controlled underwater occupancy benchmark was established, adapting representative occupancy methods to the same RGB to sonar task under a unified protocol. AquaBEV achieves 31.4 Visible IoU and 38.6 Observed IoU, 4.0% and 4.3% relative improvements over the strongest transferred baseline.