发表机构
Sun Yat-sen University; Dalian University of Technology; Southern Marine Science and Engineering Guangdong Laboratory(中山大学; 大连理工大学; 南方海洋科学与工程广东省实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对水下数据稀缺限制视觉-语言-动作模型的问题,提出免训练的代码即策略框架AquaCap,通过双层智能体、结构化感知和故障感知记忆实现自主导航与操作,仿真成功率66.43%,并在真实ROV上验证了抓取与运输能力。
AI 中文摘要
视觉-语言-动作模型的最新进展激发了人们对水下具身智能的日益关注。然而,这些模型对大规模交互数据的依赖限制了它们在水下环境中的适用性,因为在水下收集数据成本高昂且数据稀缺。为应对这一挑战,我们提出了AquaCap,一种免训练的代码即策略(Code-as-Policy)框架,用于自主水下导航与操作。AquaCap采用双层智能体架构,将任务指令和环境观测转化为条件感知的计划以及可执行的控制程序。结构化感知模块在退化的水下条件下为智能体提供语义、几何和可靠性感知的观测信息。一种故障感知记忆机制能够诊断未成功的动作,并支持闭环重规划和代码修订。这种设计使得智能体能够在线适应,而无需针对特定任务进行训练或参数更新。AquaCap在仿真中达到了66.43%的成功率。真实世界实验进一步展示了使用ROV(遥控潜水器)进行自主抓取和物体运输的能力,包括对受水动力扰动而移位的目标进行操作。
英文摘要
Recent advances in vision-language-action models have stimulated growing interest in underwater embodied intelligence. However, their reliance on large-scale interaction data limits their applicability underwater, where data collection is costly and scarce. To address this challenge, we present AquaCap, a training-free Code-as-Policy framework for autonomous underwater navigation and manipulation. AquaCap employs a dual-layer agent that translates task instructions and environmental observations into condition-aware plans and executable control programs. Structured perception then provides the agent with semantic, geometric, and reliability-aware observations under degraded underwater conditions. A failure-aware memory diagnoses unsuccessful actions and supports closed-loop replanning and code revision. This design enables online adaptation without task-specific training or parameter updates. AquaCap achieves a 66.43% success rate in simulation. Real-world experiments further demonstrate autonomous grasping and object transport with an ROV, including the manipulation of targets displaced by hydrodynamic disturbances.
Comments8 pages, 6 figures, 4 tables, 1 algorithm