发表机构
ITMO University(ITMO大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出动作条件多模态预测模型AquaJEPA,在Stonefish环境中对比多种基线,经120个带计划DVL损失的配对环境实验,其闭环性能最优,配对最终误差显著优于多数基线。
AI 中文摘要
水下机器人集成了互补传感器,其可靠性会随水体能见度、视角和载体运动发生突变。我们提出AquaJEPA,这是一种动作条件联合嵌入预测模型,融合了RGB相机、前视声呐和本体感知信息,并明确考虑传感器有效性。该模型基于8个推进器指令预测未来潜在目标,为共享的后退时域规划器提供速度和声呐剖面预测。我们在Stonefish仿真环境中对比了该方法与反应式、仅状态、普通多模态、监督式动力学及循环世界模型等基线方法,进一步分离分析了EMA目标、动作边际、掩码和模态丢弃的作用。一项预先注册的包含120个环境的复现实验,由5次独立重复实验组成,涉及3个未见过的障碍物地图、4种水体能见度系数、标称与偏移动力学,且间歇性移除DVL观测。在120个带有计划DVL损失的全新配对环境中,AquaJEPA达成74个目标,而仅状态和循环世界模型均为68个,且其平均最终误差最低,为0.906米。相较于普通多模态预测、监督式动力学和循环世界模型,配对最终误差分别降低0.273米(95%置信区间:0.190-0.356)、0.364米(0.260-0.468)和0.106米(0.025-0.187)。因此,AquaJEPA实现了最佳的整体闭环性能,在配对最终误差上显著优于3种动作条件预测基线,其相对于仅状态方法的优势在统计上仍未明确。
英文摘要
Underwater robots rely on complementary sensors whose reliability changes abruptly with water visibility and vehicle motion. We introduce AquaJEPA, a sensor-configurable family of action-conditioned joint-embedding predictive models spanning full multimodal, camera-only, sonar-only, and sensor-dropout configurations. Its members share a latent objective and receding-horizon control interface that predict future representations and physical dynamics from camera, forward-looking sonar, proprioception, and thruster commands. Trained from scratch on one hour of action-labelled data, the family is evaluated in Stonefish on 120 fresh paired scenarios spanning unseen layouts, visibility changes, dynamics shifts, and scheduled DVL loss. AquaJEPA-base achieves the strongest aggregate closed-loop performance, improving success over state-only by 12.5 percentage points and reducing final error by 0.189 m; both paired 95% intervals exclude zero. In a separate three-seed evaluation, it reduces paired final error relative to AquaJEPA-S by 0.118 m, with the same direction for every seed. AquaJEPA-robust more than halves prediction error during camera and camera-DVL blackouts. These results show that full multimodal prediction improves over state-only control and the sonar-only family member in this benchmark, while sensor-dropout training provides robustness under sensor loss.
CommentsSubmitted to IEEE ICRA 2027