发表机构
Tongji University; Tsinghua University; Zhejiang University; RoboParty Lab; Harbin Institute of Technology; ShanghaiTech University(同济大学; 清华大学; 浙江大学; RoboParty实验室; 哈尔滨工业大学; 上海科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
I-BFM首次将行为基础模型扩展至人形物体交互,通过无监督强化学习学习共享表征,单策略实现搬运、推、踢及任务链,并在跌倒后保持高成功率。
AI 中文摘要
行为基础模型(BFMs)近期已表明,单一的人形策略能够支持多样化的全身控制,但将这种通用性扩展到物理交互仍具挑战性。我们提出了I-BFM,据我们所知,这是首个用于人形物体交互的BFM。I-BFM不依赖于特定任务策略或参考轨迹跟踪,而是利用前向-反向表示和无监督强化学习,学习人形、物体及其接触之间耦合动力学的共享表征。给定下游任务奖励,同一策略可直接以潜在命令为条件执行闭环交互,无需特定任务的策略优化。为改善不同时间尺度上的交互控制,我们进一步使用短时域交互目标和长时域目标目标训练策略。单个I-BFM策略可执行搬运、推和踢,同时支持目标到达、运动跟踪、风格控制以及长时域任务链。更重要的是,在偏离名义执行的大偏差情况下,它仍然有效:在Carry任务上,I-BFM达到94.3%的名义成功率,并在机器人跌倒后保持89.3%的成功率,而基于规划基线的成功率仅为1.3%。在Unitree G1上的真实世界实验进一步展示了多样化的移动操作行为、从交互失败和外部扰动中快速恢复,以及无需特定任务重新训练的任务链执行。
英文摘要
Behavioral foundation models (BFMs) have recently shown that a single humanoid policy can support diverse whole-body control, but extending such generality to physical interaction remains challenging. We introduce I-BFM, to our knowledge the first BFM for humanoid-object interaction. Rather than relying on task-specific policies or reference tracking, I-BFM learns a shared representation of the coupled dynamics among the humanoid, objects, and their contacts using forward-backward representations and unsupervised reinforcement learning. Given a downstream task reward, the same policy can be directly conditioned on a latent command to execute closed-loop interaction without task-specific policy optimization. To improve interaction control over different time scales, we further train the policy with both short-horizon interaction targets and longer-horizon goal targets. A single I-BFM policy performs carrying, pushing, and kicking, while also supporting goal reaching, motion tracking, stylistic control, and long-horizon task chaining. More importantly, it remains effective after large deviations from nominal execution: on Carry, I-BFM achieves 94.3% nominal success and retains 89.3% success after robot falls, compared with 1.3% for a planning-based baseline. Real-world experiments on a Unitree G1 further demonstrate diverse loco-manipulation behaviors, rapid recovery from interaction failures and external disturbances, and task chaining without task-specific retraining.
Comments9 pages, including appendix