CAVE-NAV:基于VLM的水下洞穴环境自主三维导航
CAVE-NAV: VLM-Based Autonomous 3D Navigation in Underwater Cave Environments
- University of South Florida(南佛罗里达大学)
- George Mason University(乔治梅森大学)
- University of Maryland(马里兰大学)
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对水下洞穴导航的传统方法局限,提出带CoT推理的VLM导航框架,经仿真验证可完成无碰撞的端到端三维洞穴导航。
AI中文摘要:
水下洞穴环境中的自主导航对于搜救行动、科学探索和紧急撤离至关重要。传统导航系统通常依赖密集视觉特征进行定位与建图,但在水下洞穴中,视觉退化会破坏基于特征的定位,声呐建图可能生成过于保守的障碍物表示,通信限制也无法实现实时人工引导。为解决这些局限,我们提出一种自主水下洞穴导航框架,该框架利用带有思维链(CoT)推理的视觉语言模型(VLM),从环境线索(包括光强梯度、通道形态和几何复杂度)中推断可导航方向,这些线索由RGB图像、深度图和声呐垂直间隙测量组成的多模态输入捕获,从而支持在狭窄洞穴通道中实现安全的三维导航。针对多种洞穴拓扑的高保真仿真表明,所提框架完成了所有评估的端到端遍历,未发生碰撞,同时保持了与洞穴边界的安全间隙。
英文摘要:
Autonomous navigation in underwater cave environments is essential for search-and-rescue operations, scientific exploration, and emergency egress. Traditional navigation systems commonly depend on dense visual features for localization and mapping. In underwater caves, however, visual degradation can undermine feature-based localization, sonar-based mapping may yield overly conservative obstacle representations, and communication constraints preclude real-time human guidance. To address these limitations, we propose an autonomous underwater cave navigation framework that leverages a vision-language model (VLM) with Chain-of-Thought (CoT) reasoning to infer navigable directions from environmental cues, including light intensity gradients, passage morphology, and geometric complexity, captured through multimodal inputs comprising RGB imagery, depth maps, and sonar-based vertical-clearance measurements, thereby supporting safe 3D navigation through confined cave passages. High-fidelity simulations across multiple cave topologies demonstrate that the proposed framework completes all evaluated end-to-end traversals without collisions while maintaining safe clearance from cave boundaries.