arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00208cs.RO

基于交互表征与技能组合开发操作与运动技能

Developing Combined Manipulation and Locomotion Skills with Interaction Representation and Skill Composition

Fanxing Meng, Jing Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出方法使人形机器人基于发育原理学习抓取、起身等策略,通过手-物交互得分组合策略,实现抓取未见过物体零样本成功率93%、持物起身成功率96%-100%。

中文摘要 AI 辅助

本文研究如何使人形机器人基于发育原理学习运动策略,并组合这些策略以生成更复杂、实用的行为。具体而言,本文提出一种方法:(1)学习全身伸手与抓取策略;(2)将该策略与起身、行走策略组合,形成抓取、起身、行走的复杂操作与运动策略。在(1)中,方法借鉴调和分析,采用三次调和函数作为权重,通过空间卷积表征手-物空间关系;利用基于发育原理的 episode 内手指关节解耦课程,机器人可自主学习通用抓取策略,无需依赖外部数据集或预训练模型。在(2)中,方法为抓取策略与单独学习的起身策略分别提供各自的观测向量,通过手-物交互得分确定各策略控制机器人关节的时机,从而组合两种策略。实验结果显示,抓取未见过物体的零样本成功率达93%,手持物体时起身的成功率为96%-100%。研究还表明,仅当各策略在同一全身人形机器人上学习时,组合不同策略才有效,即便某策略(如运动策略)看似无需用到所有身体部位(如手指)。

英文摘要

This paper addresses how to enable a humanoid robot to learn motion policies based on developmental principles and combine policies to create more sophisticated and useful behaviors. Specifically, we present an approach to (1) learning a whole-body reaching and grasping policy and (2) combining it and a standing-up and walking policy to compose a more complex policy of manipulation and locomotion: grasping, standing up, and walking. In (1), our method draws inspiration from harmonic analysis and adopts cubic harmonics as weights to represent the hand-object spatial relationship via spatial convolution. Utilizing an intra-episode finger joint decoupling curriculum based on developmental principles, a robot can autonomously learn a generalizable grasping policy without relying on external datasets or pretrained models. In (2), our method combines the grasping policy with a separately learned getting-up policy by providing both policies with their respective observation vectors and using hand-object interaction scores to determine when each policy should control which robot joints. Our results show a 93% zero-shot success rate for grasping unseen objects and a 96-100% success rate for standing up while holding the object. Our work also demonstrates that combining different policies is only effective if each policy learning happens on the same whole humanoid body even if a policy (such as for locomotion) does not seem to need all the body parts (such as fingers).

发表机构

  • Worcester Polytechnic Institute(伍斯特理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑