WasserMan:水下操作策略学习基准
WasserMan: Benchmark for Underwater Manipulation Policy Learning
- Central University(中央大学)
- Center for Engineering Systems and Sciences(工程系统与科学中心)
- Sirius University of Science and Technology(天狼星科技大学)
- V. A. Trapeznikov Institute of Control Sciences of Russian Academy of Sciences(俄罗斯科学院特拉佩兹尼科夫控制科学研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
WasserMan是首个水下浮动基座接触操作的多任务仿真基准,提供十个任务,评估多种策略,并区分专家可行性与学习完成度。
AI中文摘要:
水下操作涉及视觉决策、接触力以及由推进器控制的浮动基座。我们提出了WasserMan,据我们所知,这是首个用于浮动基座水下接触操作的视觉运动学习的多任务仿真基准。它提供了十个专家可解决的任务、两种车臂平台、一种双臂配置以及水下动力学。其中九个任务具有学习策略评估。我们在六个任务上比较了ACT、扩散策略(DP)和行为克隆,进行了三次训练运行并采用相同的采样窗口预算。预训练的SmolVLA增加了对三个任务的评估。在测试的设置中,移除积分作用可能会阻止任务完成,而改变动作接口可能会降低学习策略的成功率,尽管专家回放成功。水流会导致任务完成和驱动力方面的任务依赖性变化。版本化的任务、演示和每集证据支持一种协议,该协议区分了专家可行性、学习完成度和执行努力。
英文摘要:
Underwater manipulation couples visual decisions, contact forces and a thruster-controlled floating base. We introduce WasserMan, to our knowledge the first multi-task simulation benchmark for visuomotor learning of floating-base underwater contact manipulation. It provides ten expert-solvable tasks, two vehicle-arm platforms, a bimanual configuration and underwater dynamics. Nine tasks have learned-policy evaluations. We compare ACT, diffusion policies (DP) and behavioral cloning on six tasks, with three training runs and equal sampled-window budgets. Pretrained SmolVLA adds evaluations on three tasks. In the tested settings, removing integral action can prevent completion, while changing action interfaces can reduce learned-policy success despite successful expert replay. Currents produce task-dependent changes in completion and actuation effort. Versioned tasks, demonstrations and per-episode evidence support a protocol separating expert feasibility, learned completion and execution effort.