将视觉基础模型适配到声学领域以实现无姿态三维声呐重建
Adapting Vision Foundation Models to Acoustics for Pose-Free 3D Sonar Reconstruction
- University of Maryland, College Park(马里兰大学帕克分校)
- Carnegie Mellon University(卡内基梅隆大学)
- University of Rhode Island(罗德岛大学)
- Dartmouth College(达特茅斯学院)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文通过利用几何关系与物理噪声模型合成数据,将视觉基础模型适配至声呐领域,首次实现无姿态三维声呐重建。
中文摘要 AI 辅助
在互联网规模RGB数据集上训练的视觉基础模型,在从文本到视频生成到少样本三维场景重建等一系列任务中展现出卓越能力。在大规模声呐数据集上训练的声学基础模型,可以在浑浊和低能见度条件下实现类似能力,而传统RGB基础模型在此类水下环境中不适用。遗憾的是,缺乏免费可用的大规模声呐数据集,使得从头训练这样的模型不切实际。在这项工作中,我们证明了视觉基础模型可以通过(1)利用两种传感模态之间的几何关系,以及(2)采用基于物理的精确噪声模型进行合成数据生成,来高效适配到声呐场景。由此产生的声呐适配模型实现了新能力:我们首次通过实验证明了基于声呐的无姿态三维重建。
英文摘要
Vision foundation models trained on Internet-scale RGB datasets enable remarkable capabilities across a range of tasks, from text-to-video generation to few-shot 3D scene reconstruction. An acoustic foundation model trained on large-scale sonar datasets could enable similar capabilities in the underwater domain, where turbidity and low-visibility conditions make conventional RGB foundation models inapplicable. Unfortunately, a lack of freely available large-scale sonar datasets makes training such a model from scratch impractical. In this work, we demonstrate that vision foundation models can be efficiently adapted to the sonar setting by (1) exploiting the geometric relationship between the two sensing modalities and (2) employing accurate physics-based noise models for synthetic data generation. The resulting sonar adaptation models enable new capabilities: For the first time, we experimentally demonstrate sonar-based pose-free 3D reconstruction.