发表机构
Tongji University; Shanghai Innovation Institute; Fudan University(同济大学; 上海创新研究院; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
X-WBC提出跨具身基础框架,分离共享人体运动语义与本体特定执行,通过因果Transformer和轻量模块实现多机器人联合训练,在模拟与真实机器人上验证了性能提升与泛化能力。
AI 中文摘要
将人形全身控制扩展到通用部署需要大规模的人体运动语料库以及跨机器人本体共享的训练经验。现有方法通常为每个机器人单独训练一个策略,导致运动经验在不同本体之间相互隔离。我们提出了X-WBC,一个跨具身基础框架,将相对共享的人体运动语义与特定本体的物理执行分离。以人为中心的命令令牌对齐完整的人体运动、机器人参考运动和稀疏的VR观测。一个因果Transformer从混合的多机器人轨迹中学习可复用的时间结构,而轻量级的机器人特定模块将共享表示映射到每个机器人的本体感觉和动作空间。在九个模拟本体、外部运动以及四个真实机器人上的实验表明,联合训练改善了跟踪性能,对齐的表示支持跨命令源的一致控制,且学习到的策略在训练语料库之外仍具有竞争力。这些结果支持异构人形机器人作为联合数据源,并确立了跨具身联合训练作为实现全身控制基础模型的实用途径。
英文摘要
Scaling humanoid whole-body control toward general-purpose deployment requires large human motion corpora and training experience shared across robot bodies. Existing methods usually train one policy per robot, leaving motion experience isolated across embodiments. We introduce X-WBC, a cross-embodiment foundation framework that separates relatively shared human motion semantics from embodiment-specific physical execution. Human-centered command tokens align full human motion, robot reference motion, and sparse VR observations. A causal Transformer learns reusable temporal structure from mixed multi-robot rollouts, while lightweight robot-specific modules map the shared representation to each robot's proprioception and action space. Across nine simulated embodiments, external motions, and four real robots, experiments show that joint training improves tracking, the aligned representation supports consistent control across command sources, and the learned policy remains competitive beyond the training corpus. These results support heterogeneous humanoids as joint data sources and establish cross-embodiment joint training as a practical route toward whole-body control foundation models.
CommentsAccepted to CoRL 2026