arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05365cs.RO

部分可观测条件下用于自主水下航行器(UUV)鲁棒导航的统一规划-学习框架

Unified Planning-Learning Framework for Robust UUV Navigation Under Partial Observability

  • NTNU(挪威科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Md Ether Deowan, Eleni Kelasidi

AI总结:

该研究提出统一规划-学习框架,结合占用地图构建、规划与控制,通过行为树蒸馏等技术提升UUV在部分可观测动态水下环境的导航鲁棒性与安全性,经仿真验证优于基线方法。

AI中文摘要:

本文提出了一种仅依赖观测的自主框架,用于动态水下环境中的自主水下航行器(UUV)导航,该框架集成了持久占用地图构建、全局避障规划和风险感知本地控制。所提流程仅利用机载声呐和深度图像观测构建占用地图,适配避障约束的全局规划器(GP)以提供长程结构,并集成强化学习(RL)策略处理短程跟踪与反应式避障。为进一步支持部分可观测下的决策,系统从机载传感器数据中学习紧凑的隐状态表示,编码环境结构、障碍物动力学与不确定性。引入带阶段性监督的行为树(BT)蒸馏以提升安全性与训练稳定性,同时采用不确定性校准的蒸馏机制,利用在线隐式模型不确定性对教师指导进行重加权,在学习过程中强调不确定区域,碰撞时间(TTC)与避障距离线索在规划和本地策略特征中保持明确。为验证框架效能,在NVIDIA Isaac Sim的高保真GPU加速仿真中建立了可复现的多随机种子评估协议,并将性能与仅BT及标准RL基线进行基准测试。所得结果表明,该框架在动态条件下提升了鲁棒性与安全性,从而提供了一种具有统一混合规划-学习架构的通用流程,以及用于部分可观测下UUV鲁棒自主的可复现方法。

英文摘要:

This paper presents an observation-only autonomy framework for Unmanned Underwater Vehicles (UUVs) navigation in dynamic underwater environments that integrates persistent occupancy mapping, global clearance-aware planning, and risk-aware local control. The proposed pipeline constructs occupancy maps solely from onboard sonar and depth image observations, adapts a clearance-constrained global planner (GP) to provide long-horizon structure, and integrates a reinforcement learning (RL) policy to handle short-range tracking and reactive avoidance. To further support decision-making under partial observability, the system learns a compact latent state representation from onboard sensor data, encoding environmental structure, obstacle dynamics, and uncertainty. Behavior tree (BT) distillation with staged supervision is introduced to improve safety and training stability, while an uncertainty-calibrated distillation mechanism reweights teacher guidance using online latent-model uncertainty, emphasizing uncertain regimes during learning, with time-to-collision (TTC) and clearance cues remaining explicit in planning and local policy features. To demonstrate the efficacy of the framework, a reproducible multi-seed evaluation protocol is established in high-fidelity GPU-accelerated simulation using NVIDIA Isaac Sim, and performance is benchmarked against BT-only and standard RL baselines. The results obtained demonstrate improved robustness and safety under dynamic conditions, thus providing a general pipeline with a unified hybrid planning learning architecture and a reproducible methodology for robust UUV autonomy under partial observability.

补充信息

↑