水下C3-JEPA:面向ROV打捞的以对象为中心的跨视角世界模型
Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage
- Underwater Engineering Institute, Shanghai Jiao Tong University(上海交通大学水下工程研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出水下C3-JEPA,一种以对象为中心的多视角预测世界模型,用于ROV打捞,通过跨视角注意力预测任务对象状态,支持MPC评估,并在真实视频中验证了其迁移能力。
AI中文摘要:
我们提出了Underwater C$^{3}$-JEPA(跨视角、控制条件、上下文扩展),这是一种面向近场重载水下ROV打捞的以对象为中心的多视角预测性世界模型。在没有接触传感器的情况下,它从同步的多视角RGB观测和车辆控制信号中,在潜在空间中预测任务对象状态如何通过接触交互以及在车辆流体动力滞后下演变。C$^{3}$-JEPA将多相机观测编码为任务对象和上下文标记,通过保留视图注意力融合跨相机证据,并直接预测以控制为条件的未来状态。弱绑定以低标注成本锚定目标和夹持器,而SIGReg则锐化了几何表示。实验表明,与无重建的潜在基线相比,学习到的表示向下游探针传递了显著更多的任务相关信息,同时保持了预测器的轻量化。由此产生的预测接口支持模型预测控制(MPC)候选评估和想象滚动行为智能体训练。在真实水下视频上的验证表明,相同的架构能够恢复被遮挡相机的对象状态,并领先于持续性预测,因此该方法可迁移到仿真之外。
英文摘要:
We present Underwater C$^{3}$-JEPA (cross-view, control-conditioned, context-extended), an object-centric multi-view predictive world model for near-field heavy-load underwater ROV salvage. Without contact sensors, it predicts in latent space how the task-object state evolves through contact interaction and under the hydrodynamic lag of the vehicle, from synchronized multi-view RGB observations and vehicle control signals. C$^{3}$-JEPA encodes multi-camera observations into task-object and context tokens, fuses cross-camera evidence through held-out-view attention, and directly predicts future states conditioned on control. Weak binding anchors the target and gripper at low annotation cost, while SIGReg sharpens the geometric representation. Experiments show that the learned representation transfers substantially more task-relevant information to downstream probes than a reconstruction-free latent baseline, while keeping the predictor lightweight. The resulting predictive interface supports model-predictive-control (MPC) candidate evaluation and imagined-rollout behavior-agent training. Validation on real underwater video shows the same architecture recovering a withheld camera's object state and staying ahead of persistence, so the recipe transfers beyond simulation.