arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkelWAM:一种用于零样本跨本体操作(Cross-Embodiment Manipulation)的骨架引导的世界-动作模型(World-Action Model)

SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation

Pengjun Niu, Yujia Xie, Rui Peng, Hang Zhao, Ke Liu

arXiv 2609.21983首次发表:更新:

发表机构

School of Advanced Manufacturing and Robotics, Peking University; Feagine; IIIS, Tsinghua University; Research Center for Robotics, Peking University(北京大学先进制造与机器人学院; Feagine; 清华大学智能产业研究院; 北京大学机器人研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SkelWAM提出骨架引导的世界-动作模型,通过共享25维状态实现零样本跨本体操作,在LIBERO-Cross10基准上以43.3%成功率超越基线36.2个百分点。

AI 中文摘要

跨机器人本体复用操作经验对于扩展机器人学习和减少重复的特定任务数据收集至关重要。然而,本体的变化会改变视觉外观、动作维度和语义,以及能够实现相同工具位姿的全身配置。我们提出了SkelWAM,一种骨架引导的世界-动作模型,通过一个显式的几何表示将感知和控制耦合起来,用于单源跨本体操作。手臂中心线几何、工具中心点(TCP)位姿和平行夹爪命令构成一个共享的25维状态。相同的定义也用于规范第三视角和腕部观测以及未来的全身动作目标。通过预测性视觉监督训练,一个视频-动作混合变换器(video-action mixture of transformers)预测规范骨架动作块,这些动作块由特定本体的约束解码器转换为关节或连续体机器人控制。该公式不需要一对一的关节对应关系,也不使用目标任务演示或目标策略更新。我们引入了LIBERO-Cross10,一个仅源域的跨本体迁移基准,涵盖十个任务和十个目标本体,跨越四个形态组。在该基准上,Franka训练的SkelWAM在1000个回合中实现了43.3%的成功率,比评估的最佳基线高出36.2个百分点。我们进一步将JAKA mini2训练的策略部署到Feagine A03连续体机器人上,用于三个桌面操作任务,展示了该方法在真实世界跨本体操作中的潜力。项目页面:this http URL

英文摘要

Reusing manipulation experience across robot embodiments is important for scaling robot learning and reducing repeated task-specific data collection. However, changes in embodiment alter visual appearance, action dimensionality and semantics, and the whole-body configurations that can realize the same tool pose. We present SkelWAM, a skeleton-guided world-action model that couples perception and control through one explicit geometric representation for single-source cross-embodiment manipulation. Arm centerline geometry, tool-center-point (TCP) pose, and parallel-jaw commands form a shared 25-D state. The same definition underlies canonical third-person and wrist observations and future whole-body action targets. Trained with predictive visual supervision, a video-action mixture of transformers predicts canonical skeleton action chunks, which embodiment-specific constrained decoders convert into joint or continuum-robot controls. This formulation requires no one-to-one joint correspondence and uses no target-task demonstrations or target policy updates. We introduce LIBERO-Cross10, a source-only cross-embodiment transfer benchmark covering ten tasks and ten target embodiments across four morphological groups. On this benchmark, Franka-trained SkelWAM achieves 43.3% success over 1,000 episodes, exceeding the best-performing evaluated baseline by 36.2 percentage points. We further deploy a JAKA mini2-trained policy on the Feagine A03 continuum robot for three tabletop manipulation tasks, illustrating the approach's potential for real-world cross-embodiment manipulation. Project page: http://www.liukepku.com/skelwam/index.html

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑