arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03234cs.RO

用于人形机器人控制的上下文感知运动先验学习

Learning Context-Aware Motion Priors for Humanoid Control

Yunyang Mo, Yi Gu, Yangchen Zhou, Hanyang Cao, Renjing Xu

首次发表
浏览论文内容

中文总结 AI 辅助

提出CMP框架,适配通用运动先验到任务上下文,在五项人形机器人控制任务中提升了性能、样本效率与鲁棒性。

中文摘要 AI 辅助

运动先验为学习自然的人形机器人行为提供了有力指导。然而,现有方法通常从整个参考数据集中学习通用的、与任务无关的先验,并在策略训练中统一应用。因此,该先验无法区分哪些参考运动与当前任务上下文相关,可能提供不相关或冲突的指导。我们提出上下文感知运动先验(Context-Aware Motion Priors, CMP),这是一种无需手动技能标签、数据集划分或单独技能发现阶段,即可将通用运动先验适配到当前任务上下文的框架。具体而言,CMP利用高优势策略rollout学习上下文-运动兼容性,同时基于演示的目标使所学相关性保持在参考分布范围内。所得相关性分数对参考监督进行重加权,以训练轻量级的上下文条件适配器。为评估CMP的有效性和通用性,我们将其分别实例化为对抗性运动先验(Adversarial Motion Priors)和评分匹配运动先验(Score-Matching Motion Priors)。在五项人形机器人控制任务中,CMP始终提升任务性能和样本效率,学习到有意义的上下文-运动对齐,且对不平衡的参考分布保持鲁棒性。这些结果表明,将运动先验适配到任务上下文为人形机器人策略学习提供了更相关的指导。

英文摘要

Motion priors provide powerful guidance for learning naturalistic humanoid behaviors. However, existing methods typically learn a general, task-agnostic prior from the entire reference dataset and apply it uniformly throughout policy training. As a result, the prior cannot distinguish which reference motions are relevant to the current task context, potentially providing irrelevant or conflicting guidance. We present Context-Aware Motion Priors (CMP), a framework that adapts a general motion prior to the current task context without manual skill labels, dataset partitioning, or a separate skill discovery stage. Specifically, CMP learns context-motion compatibility using high-advantage policy rollouts, while a demonstration-based objective keeps the learned relevance grounded in the reference distribution. The resulting relevance scores reweight reference supervision for training a lightweight context-conditioned adapter. To evaluate the effectiveness and generality of CMP, we instantiate it with both Adversarial Motion Priors and Score-Matching Motion Priors. Across five humanoid control tasks, CMP consistently improves task performance and sample efficiency, learns meaningful context-motion alignment, and remains robust to imbalanced reference distributions. These results show that adapting motion priors to task contexts provides more relevant guidance for humanoid policy learning.

补充信息

↑