发表机构
VinRobotics; National University of Singapore; TU Darmstadt; VinUniversity; DFKI; Hessian AI(VinRobotics; 新加坡国立大学; 达姆施塔特工业大学; VinUniversity; 德国人工智能研究中心; 黑森人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CrossBFM提出无机器人特定参数的统一编码器,在不到1个GPU小时内蒸馏跨人形形态的共享潜在行为空间,实现运动跟踪、目标到达和奖励优化三种提示模式的高效迁移与泛化。
AI 中文摘要
行为基础模型(BFM)赋予人形机器人在潜在行为空间上的可提示策略,使得单个向量即可表示要模仿的运动、要达到的姿态或要最大化的奖励。前向-后向表示成功生成了此类空间,但代价是单个机器人需耗费数百GPU小时。此外,当针对第二个机器人重复训练过程时,会生成与第一个无关的第二个空间,导致特定于形态的潜在表示无法统一或迁移。我们通过CrossBFM解决这些问题,将潜在空间视为各种形态的可迁移资产。由于重定向提供了帧级跨形态对应关系,我们提出了一种无机器人特定参数的统一编码器架构,用于蒸馏行为空间,从而在不到一个GPU小时内同时处理所有训练形态。遵循该编码器,潜在条件跟踪器以传统PPO训练方式在仅额外10个GPU小时内将蒸馏后的潜在表示转化为全身控制。在三个蒸馏人形机器人上,所有三种提示模式均可迁移:潜在条件策略的运动跟踪仅比其关节条件对应物损失0.025弧度,姿态间平滑目标到达无摔倒,以及所有41个奖励提示的奖励优化。我们的实验进一步揭示:1)在四分之一运动语料上回归编码器仅损失5%的跟踪性能;2)在机器人子集上训练编码器并在未见机器人上评估,可恢复所见机器人跟踪性能的89%,展示了形态相似机器人的跨形态泛化能力。我们还在真实机器人上验证了该流程,涵盖所有三种提示模式及基于流生成的潜在表示。项目网站:此https URL
英文摘要
Behavior Foundation Models (BFMs) give humanoids a promptable policy over a latent behavior space, enabling one single vector to represent a motion to imitate, a pose to reach, or a reward to maximize. Forward-Backward representations successfully produce such spaces, but at the cost of hundreds of GPU-hours for a single robot. Moreover, when the training process is repeated for a second robot, it produces a second space unrelated to the first, resulting in embodiment-specific latents that do not unify or transfer. We address these problems with CrossBFM, treating the latent space as the transferable asset for various embodiments. As retargeting provides frame-level cross-embodiment correspondence, we propose a unified encoder architecture with no robot-specific parameters for distilling the behavior space to address all training embodiments simultaneously in less than a GPU-hour. Following this encoder, latent-conditioned trackers turn the distilled latent into whole-body control in a conventional PPO training manner in just 10 more GPU-hours. On three distilled humanoids, all three prompting modes transfer: motion tracking with latent-conditioned policy losing only $0.025$ rad to its joint-conditioned counterpart, smooth goal reaching between poses with no falls, and reward optimization for all $41$ reward prompts. Our experiments further reveal that 1) regressing the encoder on a quarter of the motion corpus costs only $5\%$ of tracking performance and 2) training the encoder on a subset of robots and evaluating on an unseen one recovers up to $89\%$ of the tracking performance of seen robots, demonstrating cross-embodiment generalization to morphologically similar robots. We also verify the pipeline on real robots across all three prompting modes and with flow-based generated latents. Project website: https://dotandung.github.io/crossbfm/
CommentsProject Website: https://dotandung.github.io/crossbfm/