arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MASkillBlender:通过技能混合实现多人类形机器人去中心化全身协调操作

MASkillBlender: Decentralized Whole-Body Coordination for Multi-Humanoid Loco-Manipulation via Skill Blending

Yifan Hu, Luhang Hong, Mingkang Long, Danning Wang, Chengfeng Jia, Rong Su, Junjie Fu, Guanghui Wen

arXiv 2610.01102首次发表:更新:

发表机构

Nanyang Technological University; Southeast University; Purple Mountain Laboratories(南洋理工大学; 东南大学; 紫金山实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MASkillBlender,一种通用多智能体强化学习框架,通过共享去中心化高层策略混合预训练技能,仅用任务级奖励实现多人类形机器人去中心化全身协调,并引入排列数据增强提升效率,仿真验证有效。

AI 中文摘要

协调的多人类形机器人移动操作任务前景广阔,但由于高维全身控制、去中心化决策和可扩展性等问题,实现起来颇具挑战。尽管近期的强化学习方法已改善了单个人类形机器人的全身控制,但将其扩展到多人类形机器人场景仍非易事,且通常需要大量的奖励工程或任务特定设计。我们提出了MASkillBlender,一个通用的多智能体强化学习框架,用于实现去中心化的多人类形机器人全身协调。通过学习一个共享的、基于可复用预训练单人类形机器人技能的去中心化高层策略,MASkillBlender仅使用任务级奖励即可实现协调行为,无需任务特定的运动参考。为提高学习效率,我们进一步引入了一种基于排列的数据增强策略,适用于同构多人类形机器人系统,并从理论上证明,在同构马尔可夫博弈公式下,排列后的样本保持了原始样本的策略梯度方向。我们在两种人类形机器人实体上对多个多人类形机器人协调任务进行了评估。仿真结果表明,所提出的框架在不同任务和人类形机器人实体上均能持续实现强大的任务性能并产生协调行为。

英文摘要

Coordinated multi-humanoid loco-manipulation is promising yet challenging due to high-dimensional whole-body control, decentralized decision making, and scalability. While recent reinforcement learning methods have improved single-humanoid whole-body control, extending them to the multi-humanoid setting remains nontrivial and often requires substantial reward engineering or task-specific design. We propose MASkillBlender, a general multi-agent reinforcement learning framework to achieve decentralized multi-humanoid whole-body coordination. By learning a shared decentralized high-level policy over reusable pre-trained single-humanoid skills, MASkillBlender enables coordinated behaviors using only task-level rewards, without requiring task-specific motion references. To improve learning efficiency, we further introduce a permutation-based data augmentation strategy for homogeneous multi-humanoid systems, and theoretically show that the permuted samples preserve the policy-gradient direction of the original samples under the homogeneous Markov game formulation. We evaluate MASkillBlender on multiple multi-humanoid coordination tasks across two humanoid embodiments. Simulation results demonstrate that the proposed framework consistently achieves strong task performance and enables coordinated behaviors across different tasks and humanoid embodiments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑