AI 中文总结
本文提出AquaWAM,首个面向水下具身智能体的世界动作模型,建模动作与被动动力学,在USIM基准上以72.6%成功率超越现有方法,决策速度快2.7倍。
AI 中文摘要
世界动作模型(WAMs)在具身智能领域正变得日益重要和有用,它们使机器人能够在与物理环境交互之前预判候选动作的后果。然而,水下机器人通常受到被动动力学的影响,如惯性、浮力、水动力阻力和持续漂移,这些效应即使在动作完成后仍可能继续影响载具。现有的WAMs主要预测动作条件下的视觉观测,并未明确设计用于捕捉此类被动运动动力学。本文提出AquaWAM,这是首个专为水下具身智能体设计的世界动作模型。AquaWAM不预测未来图像,而是同时建模动作条件动力学和被动物理动力学,包括推进器死区、超出每条指令持续时间的惯性滑行以及环境流。具体而言,它通过多普勒测速仪(DVL)、惯性测量单元(IMU)、压力传感器和关节编码器进行感知,而相机仅提供用于理解目标和目标位姿的语义信息。通过建模紧凑的导航状态而非高维视觉观测,AquaWAM相比传统WAMs大幅降低了模型尺寸和计算成本。实验方面,AquaWAM在USIM基准的20项水下任务中达到72.6%的任务成功率,优于现有方法,同时在NVIDIA Jetson AGX Orin上的动作决策速度比U0快2.7倍。此外,当部分机载传感器测量不可用时,模型仍保持有效性。例如,在没有DVL速度测量的情况下,我们的方法仍达到61.6%的成功率,而U0仅为39.4%。
英文摘要
World Action Models (WAMs) are becoming increasingly important and useful for embodied intelligence, as they enable robots to anticipate the consequences of candidate actions before interacting with the physical environment. However, underwater robots are usually subject to passive dynamics, such as inertia, buoyancy, hydrodynamic drag, and persistent drift, which can continue to affect the vehicle even after an action is completed. Existing WAMs, which primarily predict action-conditioned visual observations, are not explicitly designed to capture such passive motion dynamics. In this paper, we present AquaWAM, the first World Action Model designed for underwater embodied agents. Instead of predicting future images, AquaWAM models both action-conditioned and passive physical dynamics, including the thruster dead band, the inertial glide that outlasts each command, and ambient currents. Specifically, it senses through the DVL, IMU, pressure sensor and joint encoders, while cameras supply only semantics for understanding goals and target pose. By modeling compact navigation states rather than high-dimensional visual observations, AquaWAM substantially reduces the model size and computational cost compared with conventional WAMs. Experimentally, AquaWAM achieves a 72.6% task success rate across 20 underwater tasks on the USIM benchmark, outperforming existing methods while making action decisions 2.7x faster than U0 on an NVIDIA Jetson AGX Orin. Our model also remains effective when some onboard sensor measurements are unavailable. For example, without DVL velocity measurements, our method still achieves a 61.6% success rate, compared with 39.4% for U0.