RoboStriker:用于自主人形机器人拳击的隐空间策略博弈
RoboStriker: Latent-Space Strategic Games for Autonomous Humanoid Boxing
- Shanghai Jiao Tong University(上海交通大学)
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- Shanghai Innovation Institute(上海创新研究院)
- Huazhong University of Science and Technology(华中科技大学)
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究提出RoboStriker框架,将人形拳击任务形式化为隐空间零和马尔可夫博弈,通过分层结构解耦推理与执行,实现了优于原始动作空间方法的战术性能并成功部署到现实人形机器人。
中文摘要 AI 辅助
在人形机器人中实现人类水平的竞争智能与身体敏捷性仍是一项重大挑战,尤其是在拳击这类接触密集且高度动态的任务中。尽管多智能体强化学习为策略交互提供了原则性框架,但将其直接应用于非结构化原始运动空间不可避免地会导致关节级物理崩溃,阻碍任何可行战斗战术的出现。为解决策略探索与物理可行性之间的这一根本矛盾,我们将人形战斗任务形式化为一种新型双人隐空间零和马尔可夫博弈。在标准正则性和近似最优响应假设下,我们证明该隐空间形式化在解码器可达动作流形上诱导出等价博弈,为产生的自博弈动力学提供近似纳什均衡解释。为实例化该理论形式化,我们提出RoboStriker,这是一种将高层推理与低层执行解耦的分层框架。它首先将预定义拳击动作的跟踪专长提炼为拓扑有界的隐流形。该结构化隐基础随后通过隐空间神经虚构自我博弈驱动多智能体协同进化。大量实验结果表明,在该结构化隐空间内博弈显著优于直接探索。通过预训练动作解码器约束策略探索,RoboStriker大幅降低了原始动作空间方法中观察到的灾难性平衡失败,并在竞争胜率和打击效率两方面实现了更优的战术性能。最后,我们成功将学习到的战斗策略部署到现实世界人形机器人上并进行了验证。我们的代码、视频及补充材料可在RoboStriker获取。
英文摘要
Achieving human-level competitive intelligence and physical agility in humanoid robots remains a profound challenge, particularly in contact-rich and highly dynamic tasks such as boxing. While Multi-Agent Reinforcement Learning offers a principled framework for strategic interaction, its direct application to unstructured raw motor spaces inevitably leads to joint-level physical collapse, preventing the emergence of any viable combat tactics. To resolve this fundamental conflict between strategic exploration and physical feasibility, we formulate the humanoid combat task as a novel two-player latent-space zero-sum Markov game. Under standard regularity and approximate best-response assumptions, we show that the latent formulation induces an equivalent game over the decoder-reachable action manifold, providing an approximate-Nash interpretation of the resulting self-play dynamics. To instantiate this theoretical formulation, we propose RoboStriker, a hierarchical framework that decouples high-level reasoning from low-level execution. It first distills the tracking expertise of predefined boxing motions into a topologically bounded latent manifold. This structured latent foundation subsequently drives multi-agent co-evolution via Latent-Space Neural Fictitious Self-Play. Extensive experimental results demonstrate that gaming within this structured latent space substantially outperforms direct exploration. By constraining strategic exploration through a pretrained motion decoder, RoboStriker substantially reduces the catastrophic balance failures observed in raw action-space methods and achieves superior tactical performance in both competitive win rates and striking efficiency. Finally, we successfully deploy and validate our learned combat policies on real-world humanoid robots. Our code and video and supplementary materials are available at RoboStriker.