发表机构
ETH Zürich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出分层框架KINO,以运动关键帧连接VLM规划与RL控制,通过显著性采样策略将人形移动操作成功率从44%提升至92%,并在仿真和Unitree G1上验证。
AI 中文摘要
人形机器人的移动操作要求机器人在执行协调的全身运动时,能够理解任务指令和场景语义。我们提出了一种分层框架,该框架使用运动关键帧作为视觉语言模型(VLM)规划与强化学习(RL)控制之间的中间表示。每个关键帧指定一个目标全身机器人姿态,并在适用时指定物体姿态。给定语言指令、场景观察和执行反馈,VLM从预定义的关键帧库中选择连续的任务相关关键帧。所选关键帧被重新定向到当前场景,以考虑物体姿态和尺寸。随后,一个由关键帧条件化的全身策略生成关节级动作以达到这些目标。我们引入了一种基于显著性的关键帧采样策略用于低级策略训练,在使用稀疏VLM关键帧时,将端到端任务成功率从44%提升至92%。我们在仿真环境和Unitree G1人形机器人上对物体拾取、搬运和放置任务进行了框架评估。该系统成功执行了单手和双手操作,并能够泛化到超出训练参考数据的放置位置。
英文摘要
Humanoid loco-manipulation requires robots to interpret task instructions and scene semantics while executing coordinated whole-body motions. We propose a hierarchical framework that uses motion keyframes as an intermediate representation between Vision-Language Model (VLM) planning and Reinforcement Learning (RL) control. Each keyframe specifies a target whole-body robot pose and, when applicable, an object pose. Given a language instruction, scene observations, and execution feedback, the VLM selects successive task-relevant keyframes from a predefined library. The selected keyframes are retargeted to the current scene to account for object poses and dimensions. A keyframe-conditioned whole-body policy then generates joint-level actions to reach these goals. We introduce a saliency-based keyframe sampling strategy for low-level policy training that improves end-to-end task success rate from 44% to 92% when using sparse VLM keyframes. We evaluate our framework on object pickup, transport, and placement tasks in simulation and on a Unitree G1 humanoid. The system successfully performs both one- and two-handed manipulation and generalises to placement locations beyond the training reference data.