发表机构
Horizon Robotics(地平线机器人)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Holo-M,首个离散VLA模型,通过统一动作标记器分解全身动作空间并扩展语言模型词汇,结合分组离散扩散解码,在SIMPLE基准上实现最高成功率。
AI 中文摘要
使用离散动作标记的视觉-语言-动作(VLA)模型已被证明在机械臂操控任务中有效。然而,对于人形机器人,全身动作空间——包括腿部、躯干、手臂和手——维度更高且异质性更强,这带来了以往VLA模型未解决的标记化、训练和实时推理挑战。我们提出了Holo-M,据我们所知,这是首个用于人形机器人全身操控的离散VLA模型,它通过扩展语言模型的词汇表加入动作标记,从而内在地利用语言模型。在该模型中,我们设计了一个统一的动作标记器,将人形机器人的动作空间分解为四个针对身体部位的标记器——末端执行器、身体、手和运动学——从而支持在截然不同的实体和数据源上进行训练,包括人形机器人遥操作、自我中心视角的人类视频和仿真。通过将这些动作标记扩展进语言模型的词汇表,我们避免了使用独立连续动作专家的模型所固有的知识隔离问题。为了满足实时控制要求,我们通过分组离散扩散解码来解码每个身体部位的动作标记,而不是对动作标记使用自回归。我们在SIMPLE人形机器人全身操控基准上进行了大量实验,Holo-M在通用和专家评估中均取得了最高成功率,显著领先于第二名。我们将发布所有代码和模型权重。
英文摘要
Vision-language-action (VLA) models using discrete action tokens have proven effective for controling robotic arms on manipulation tasks. For a humanoid, however, the whole-body action space -- legs, torso, arms, and hands -- is far higher-dimensional and heterogeneous, raising tokenization, training, and real-time inference challenges that the previous VLA models do not address. We present Holo-M, to our knowledge the first discrete VLA model for humanoid loco-manipulation that intrinsically exploits the language model by extending its vocabulary with action tokens. In this model, we devise a unified action tokenizer that decomposes the humanoid action space into four body-part-specific tokenizers -- end-effector, body, hand, and kinematics -- enabling training across drastically different embodiments and data sources, including humanoid teleoperation, ego-centric human video, and simulation. By extending the language model's vocabulary with these action tokens, we avoid the knowledge-insulation problem inherent to the models that use separate continuous action experts. To meet real-time control requirements, we decode each body part's action tokens through grouped discrete diffusion decoding, rather than using autoregression on the action tokens. We have conducted extensive experiments on the SIMPLE humanoid loco-manipulation benchmark, in which Holo-M achieves the highest success rates in both the generalist and specialist evaluations, leading the second best by significant margins. We will release all the code and model weights.