发表机构
DeepMirror Inc.; HKUST; Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(深镜公司; 香港科技大学; 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出LooperMuscle框架,结合结构化混合专家等组件,弥合FastSAC与PPO在人形机器人全身跟踪任务中的速度-性能差距,训练效率高且跟踪精度优。
AI 中文摘要
FastSAC类方法大幅缩短人形机器人运动训练时间,但在全身跟踪任务中,与PPO相比常出现明显性能下降。本文针对该速度-性能差距,提出LooperMuscle,一种复合专家策略学习框架,在保持高训练效率的同时恢复跟踪质量。LooperMuscle结合语义结构化混合专家演员、专家感知分布评论家及带延迟课程调度的贡献路由回放。这三个组件形成闭环训练循环,专家贡献指导数据路由,路由后的数据塑造价值学习,价值梯度反过来优化专家专业化。实验表明,本文方法在运动跟踪精度上显著优于普通FastSAC,且所需墙上时间远少于PPO:FastSAC训练约15分钟但表现不佳,PPO结果更好但需约6小时,LooperMuscle在约45分钟模拟训练中弥合了与PPO的大部分差距,为快速策略迭代提供实用效率。代码将发布至该httpsURL以造福研究界。
英文摘要
FastSAC-style methods significantly reduce humanoid motion training time but often suffer from notable performance degradation compared with PPO in whole-body tracking tasks. We target this speed-performance gap by introducing LooperMuscle, a composed expert policy learning framework that restores tracking quality while preserving high training efficiency. LooperMuscle combines a semantically structured mixture-of-experts actor, an expert-aware distributional critic, and contribution-routed replay with deferred curriculum scheduling. These three components form a closed training loop in which expert contributions guide data routing, routed data shape value learning, and value gradients in turn refine expert specialization. Empirically, our approach substantially outperforms vanilla FastSAC in motion tracking accuracy while requiring far less wall-clock time than PPO: where FastSAC trains in about 15 minutes but underperforms, and PPO achieves stronger results but requires about 6 hours, LooperMuscle recovers a substantial fraction of the remaining gap to PPO in roughly 45 minutes of simulation training, delivering practical efficiency for rapid policy iteration. The code will be released to benefit the research community at https://loopermuscle.github.io/.