首个可部署动态-重心矩(dynamic-CoM):人形机器人单腿平衡的统一策略与方法无关基准
A Change of Frame Makes the Capture Point Proprioceptive: Distillation-Free Humanoid Single-Leg Balance
浏览论文内容
中文总结 AI 辅助
针对人形机器人单腿平衡的通用策略无法实现干净站立的问题,提出首个可部署动态-CoM策略FDDC,结合相关方法与基准,大幅提升平衡成功率并实现硬件迁移,推动平衡能力的标准化衡量。
中文摘要 AI 辅助
人形机器人的统一策略可处理敏捷的全身运动,但在一项简单需求上却表现不佳:单腿站立时保持平衡。在我们的单腿平衡基准测试中,8个已发布的最先进通用策略在90个测试动作中,没有一个能实现干净的单腿站立姿态;它们仅通过迈步或跳跃来维持平衡,属于失衡后的恢复而非预防。预防失衡需要捕获点(xCoM),即通过速度外推得到的质心(CoM),但该点从未用于硬件策略,因为它需要基座线速度,而机载传感器无法提供;若相对于支撑脚表达,该速度会完全抵消,从而可仅通过编码器和惯性测量单元(IMU)重构观测值。我们将这一首个可部署的动态-CoM观测值直接输入到运行在硬件上的执行器中,并搭配从人体姿势控制逐词翻译而来的奖励库,遵循“预防优先于修复”的原则。通过带特权评论家且无蒸馏的非对称FastSAC训练,得到的策略FDDC(首个可部署动态-CoM)在9个分层姿态类别的90个保留动作中,有86个实现了干净的单腿平衡,并可迁移到真实的Unitree G1机器人;消融实验显示,动态-CoM观测值是最大的单一驱动因素:仅移除它就会使干净单腿平衡得分降低40分。我们发布了完整的技术栈,包括首个与方法无关、可复现的人形机器人单腿平衡的仿真到仿真基准,在与训练仿真不同的测试仿真中对每个策略进行评分,朝着将平衡从特定任务技巧转变为该领域可衡量的能力迈出了一步。
英文摘要
Unified humanoid policies track agile whole-body motion, yet few can hold a clean single-leg stance. On the single-leg balance benchmark introduced here, eight released state-of-the-art general policies hold a clean stance on none of 90 held-out motions; they stay upright only by hopping or re-planting a foot, recovering from imbalance rather than preventing it. Prevention needs the capture point, the center of mass (CoM) extrapolated by its velocity. That velocity contains the base linear velocity, which no on-board sensor measures, so the capture point has been confined to rewards and privileged critics and has reached hardware only through teacher--student distillation. A change of frame removes the obstacle: expressed relative to the support foot, the base velocity cancels identically, leaving a capture-point state reconstructible from joint encoders and an inertial measurement unit alone. We place this support-relative dynamic-CoM observation directly in the deployed actor and pair it with a reward library translated term by term from human postural control. Trained with asymmetric FastSAC and no distillation, the resulting policy, DDC, holds clean single-leg balance on 89 of 90 held-out motions across nine pose classes and runs directly on a Unitree G1; removing the observation alone costs 43 points, and 53 under deployment noise. We release the policy, the data, and a method-agnostic MuJoCo benchmark for humanoid single-leg balance, which scores released policies on the same held-out motions. Together these turn single-leg balance from a per-task demonstration into a capability the field can measure and build into general policies.