arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

I-BFM:基于无监督强化学习的奖励条件鲁棒人形交互

I-BFM: Reward-Conditioned Robust Humanoid Interaction via Unsupervised Reinforcement Learning

Ziqi Han, Yitang Li, Junhan Sun, Fanrong Dong, Yaojie Shen, Lei Ye, Zetong Jing, Yongqi Zhang, Yiming Zhang, Xue Wang, Hao Zhao

arXiv 2610.06129首次发表:更新:

发表机构

Tongji University; Tsinghua University; Zhejiang University; RoboParty Lab; Harbin Institute of Technology; ShanghaiTech University(同济大学; 清华大学; 浙江大学; RoboParty实验室; 哈尔滨工业大学; 上海科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

I-BFM首次将行为基础模型扩展至人形物体交互,通过无监督强化学习学习共享表征,单策略实现搬运、推、踢及任务链,并在跌倒后保持高成功率。

AI 中文摘要

行为基础模型(BFMs)近期已表明,单一的人形策略能够支持多样化的全身控制,但将这种通用性扩展到物理交互仍具挑战性。我们提出了I-BFM,据我们所知,这是首个用于人形物体交互的BFM。I-BFM不依赖于特定任务策略或参考轨迹跟踪,而是利用前向-反向表示和无监督强化学习,学习人形、物体及其接触之间耦合动力学的共享表征。给定下游任务奖励,同一策略可直接以潜在命令为条件执行闭环交互,无需特定任务的策略优化。为改善不同时间尺度上的交互控制,我们进一步使用短时域交互目标和长时域目标目标训练策略。单个I-BFM策略可执行搬运、推和踢,同时支持目标到达、运动跟踪、风格控制以及长时域任务链。更重要的是,在偏离名义执行的大偏差情况下,它仍然有效:在Carry任务上,I-BFM达到94.3%的名义成功率,并在机器人跌倒后保持89.3%的成功率,而基于规划基线的成功率仅为1.3%。在Unitree G1上的真实世界实验进一步展示了多样化的移动操作行为、从交互失败和外部扰动中快速恢复,以及无需特定任务重新训练的任务链执行。

英文摘要

Behavioral foundation models (BFMs) have recently shown that a single humanoid policy can support diverse whole-body control, but extending such generality to physical interaction remains challenging. We introduce I-BFM, to our knowledge the first BFM for humanoid-object interaction. Rather than relying on task-specific policies or reference tracking, I-BFM learns a shared representation of the coupled dynamics among the humanoid, objects, and their contacts using forward-backward representations and unsupervised reinforcement learning. Given a downstream task reward, the same policy can be directly conditioned on a latent command to execute closed-loop interaction without task-specific policy optimization. To improve interaction control over different time scales, we further train the policy with both short-horizon interaction targets and longer-horizon goal targets. A single I-BFM policy performs carrying, pushing, and kicking, while also supporting goal reaching, motion tracking, stylistic control, and long-horizon task chaining. More importantly, it remains effective after large deviations from nominal execution: on Carry, I-BFM achieves 94.3% nominal success and retains 89.3% success after robot falls, compared with 1.3% for a planning-based baseline. Real-world experiments on a Unitree G1 further demonstrate diverse loco-manipulation behaviors, rapid recovery from interaction failures and external disturbances, and task chaining without task-specific retraining.

Comments9 pages, including appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑