arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Athena-WBC:用于长尾人形机器人全身控制的能力对齐策略专家

Athena-WBC: Capability-Aligned Policy Experts for Long-Tail Humanoid Whole-Body Control

Yuan Jiang, Ningyuan Zhang, Xicun Yang, Yuzhi Jiang, Jie Chen

arXiv 2607.04837首次发表:更新:

发表机构

XPENG ROBOTICS(XPENG机器人)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现仅靠重新分配训练精力改进人形运动跟踪控制器不完整。提出Athena-WBC,含能力对齐策略专家,动态专家用特定目标,平衡专家用重力课程,经蒸馏和微调,提升长尾运动恢复及跟踪性能。

AI 中文摘要

大规模人形运动跟踪控制器通常通过重新分配训练精力来改进:对困难运动增加采样、分成更小子集或分配给专门专家。我们表明这种观点不完整。在强大的全身控制基线中,即使经过有针对性的训练,仍有一组可行的训练片段未解决,尤其是高动态过渡和对平衡至关重要的运动。这些失败不仅源于曝光不足,还源于运动需求与默认训练方法所诱导的有效能力之间的不匹配。我们提出了Athena-WBC,这是一种用于长尾人形机器人全身控制的紧凑师生管道,具有能力对齐的策略专家。动态专家使用以跟踪为重点、感知约束的目标,在保留物理可行性约束的同时消除保守精力和时间控制惩罚;平衡专家使用重力课程来提高早期训练的生存能力。由此产生的特权教师通过DAgger蒸馏进行运动路由,然后压缩成一个具有可部署观测值的单一控制器,随后进行强化学习微调。在全尺寸人形机器人上的实验表明,与强大的SONIC方法基线相比,仅使用少量专家就能更好地恢复训练集长尾运动并具有更好的测试跟踪效果。

英文摘要

Large-scale humanoid motion-tracking controllers are commonly improved by reallocating training effort: difficult motions are sampled more often, isolated into smaller subsets, or assigned to specialized experts. We show that this view is incomplete. In strong whole-body-control baselines, a residual set of feasible training clips remains unsolved even under targeted training, especially for high-dynamic transitions and balance-critical motions. These failures arise not only from insufficient exposure, but from a mismatch between the motion demands and the effective capability induced by the default training recipe. We propose Athena-WBC, a compact teacher-student pipeline with capability-aligned policy experts for long-tail humanoid whole-body control. Dynamic experts use a tracking-focused, constraint-aware objective that removes conservative effort and temporal-control penalties while preserving physical feasibility constraints; balance experts use a gravity curriculum to improve early-training survivability. The resulting privileged teachers are motion-routed for DAgger distillation and then compressed into a single controller with deployable observations followed by RL fine-tuning. Experiments on a full-size humanoid show improved recovery of training-set long-tail motions and better held-out tracking than a strong SONIC-recipe baseline, using only a small number of experts.

CommentsWithdrawn by the authors due to unresolved authorization and data-governance concerns affecting the dataset used in the experiments and, consequently, the reported results. Readers should not rely on this work while these issues are under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑