arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15163cs.ROcs.AI

人形机器人的缩放行为基础模型

Scaling Behavior Foundation Model for Humanoid Robots

Weishuai Zeng, Kangning Yin, Xiaojie Niu, Shunlin Lu, Weixiang Zhong, Jiahe Chen, Feiyu Jia, Xiao Chen, Zirui Wang, Furui Xu, Ming Zhou, Kailin Li, Weinan Zhang… 展开作者

Weishuai Zeng, Kangning Yin, Xiaojie Niu, Shunlin Lu, Weixiang Zhong, Jiahe Chen, Feiyu Jia, Xiao Chen, Zirui Wang, Furui Xu, Ming Zhou, Kailin Li, Weinan Zhang, He Wang, Li Yi, Dahua Lin, Jiangmiao Pang, Jingbo Wang

首次发表
浏览论文内容

中文总结 AI 辅助

研究人形机器人行为基础模型的缩放问题,通过协调运动跟踪学习范式、策略展开数量与参考运动多样性的协同、人形Transformer模型架构这三个核心组件,显著提升控制保真度和任务泛化能力,降低测试集误差。

中文摘要 AI 辅助

人形控制需要自然的全身协调、对控制信号的精确实时响应以及在不同环境背景下的强大泛化能力,这使其成为通用具身智能体的基石。行为基础模型(BFMs)最近成为一种有前途的解决方案,通过利用大规模行为数据来应对这些挑战,以实现卓越的表现力、通用性和泛化能力。然而,尽管人们越来越关注扩展BFMs以进一步提高其能力,但尚不清楚包括学习范式、行为数据和模型架构等关键因素应如何协调以实现有效扩展。在这项工作中,我们重新审视了BFMs的扩展方法,并证明通过协调三个核心组件可以实现显著的性能提升:1)运动跟踪的学习范式,将各种人形控制问题重新表述为在全局框架中再现集成的全身行为;2)策略性地在策略展开数量和参考运动多样性之间协同;3)称为人形Transformer的具有表现力和可扩展的模型架构,促进结构化行为表示的自然出现。通过在模拟和实际部署中的广泛实验,我们证明我们的方法在控制保真度和任务泛化方面有显著改进,与现有人形控制器相比,在局部模式下测试集上的平均每关键点位置误差(MPKPE)降低了10%以上,在全局模式下降低了82%。这些结果将BFM确立为可扩展和通用人形控制的原则性和有效基础。

英文摘要

Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) have recently emerged as a promising solution to address these challenges by leveraging large-scale behavioral data to achieve superior expressiveness, versatility and generalization. However, despite growing interest in scaling BFMs to further improve their capabilities, it remains unclear how key factors, including the learning paradigm, behavioral data and model architecture should be coordinated to enable effective scaling. In this work, we revisit the scaling recipe for BFMs and demonstrate that substantial performance gains can be achieved through the coordination of three core components: 1) the learning paradigm of motion tracking that reformulates diverse humanoid control problems as the reproduction of integrated whole-body behaviors in the global frame; 2) the strategic synergy between on-policy rollout quantity and reference motion diversity; and 3) the expressive and scalable model architecture termed Humanoid Transformer that facilitates the natural emergence of structured behavioral representations. Through extensive experiments in both simulation and real-world deployment, we demonstrate that our approach yields significant improvements in control fidelity and task generalization, reducing Mean Per-Keypoint Position Error (MPKPE) on the test set by over 10% in local mode and 82% in global mode compared with existing humanoid controllers. These results establish BFM as a principled and effective foundation for scalable and general-purpose humanoid control.

↑