arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24054cs.RO

AquaOrbit:间歇性视觉反馈下水下目标环绕的仿真到现实强化学习

AquaOrbit: Sim-to-Real Reinforcement Learning for Underwater Target Orbiting under Intermittent Visual Feedback

Kanzhong Yao, Jinyi Leng, Hao Zhang, Zhe Sun, Xuelong Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对水下环绕中视觉中断问题,提出带恢复模块的强化学习控制器AquaOrbit,经仿真训练后零样本部署,在移动目标下视线误差降低46%,并成功完成多种轨迹。

中文摘要 AI 辅助

在水下环绕过程中,间歇性的视觉丢失会破坏目标相对反馈,使得维持协调运动并重新捕获移动目标变得困难。我们提出了AquaOrbit,一种带有恢复模块的强化学习控制器,用于在视觉反馈中断的情况下进行水下目标环绕。在检测丢失期间,恢复模块利用锁存视线、横滚和深度参考来支持稳定和目标重新捕获。我们在Isaac Sim中训练该控制器,并加入了动力学、观测和视觉丢失的随机化。在未重新训练的情况下,在采用不同物理引擎和感知扰动的Gazebo/ROS2中评估,AquaOrbit在静态和移动目标条件下,在未见过的变深度3D轨迹上均完成了20/20次环绕试验。在移动目标条件下,相对于带有恢复的基于PID的视觉伺服控制器,它平均将视线误差降低了约46%,同时保持了相当的路径跟踪精度;移除恢复模块后,完成率降至9/20。零样本物理部署,完全机载感知和控制,展示了椭圆、八字形和变深度圆形轨迹,包括训练中未出现的后两种轨迹类型。在持续长达8秒的人工遮挡期间,机器人保持姿态稳定,并在所报告的姿态引起的视场丢失事件中,在2.5秒内重新捕获目标。

英文摘要

Intermittent visual loss disrupts target-relative feedback during underwater orbiting, making it difficult to maintain coordinated motion and reacquire a moving target. We present AquaOrbit, a reinforcement-learning controller with a recovery module for underwater target orbiting under interrupted visual feedback. During detection loss, the recovery module uses latched line-of-sight, roll, and depth references to support stabilization and target reacquisition. We train the controller in Isaac Sim with dynamics, observation, and vision-loss randomization. Evaluated without retraining in Gazebo/ROS2 under a different physics engine and perception perturbations, AquaOrbit completes 20/20 orbiting trials in each of the static- and moving-target conditions on an unseen variable-depth 3-D trajectory. In the moving-target condition, it reduces mean line-of-sight error by approximately 46% relative to a PID-based visual servoing controller with recovery while maintaining comparable path-tracking accuracy; removing the recovery module reduces completion to 9/20. Zero-shot physical deployment with fully onboard perception and control demonstrates elliptical, figure-eight, and variable-depth circular trajectories, including the latter two trajectory types absent from training. The robot maintains attitude stability during manual occlusions lasting up to 8s and reacquires the target within 2.5s in the reported attitude-induced field-of-view loss events.

发表机构

  • Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究院(TeleAI))
  • Harbin Engineering University(哈尔滨工程大学)
  • Tongji University(同济大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑