arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KINO:面向人形机器人移动操作中VLM规划与全身控制的关键帧接口

KINO: A Keyframe Interface for VLM Planning and Whole-Body Control in Humanoid Loco-Manipulation

Sitong Chen, Fatemeh Zargarbashi, Jin Cheng, Tianxu An, Stelian Coros

arXiv 2609.18869首次发表:更新:

发表机构

ETH Zürich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出分层框架KINO,以运动关键帧连接VLM规划与RL控制,通过显著性采样策略将人形移动操作成功率从44%提升至92%,并在仿真和Unitree G1上验证。

AI 中文摘要

人形机器人的移动操作要求机器人在执行协调的全身运动时,能够理解任务指令和场景语义。我们提出了一种分层框架,该框架使用运动关键帧作为视觉语言模型(VLM)规划与强化学习(RL)控制之间的中间表示。每个关键帧指定一个目标全身机器人姿态,并在适用时指定物体姿态。给定语言指令、场景观察和执行反馈,VLM从预定义的关键帧库中选择连续的任务相关关键帧。所选关键帧被重新定向到当前场景,以考虑物体姿态和尺寸。随后,一个由关键帧条件化的全身策略生成关节级动作以达到这些目标。我们引入了一种基于显著性的关键帧采样策略用于低级策略训练,在使用稀疏VLM关键帧时,将端到端任务成功率从44%提升至92%。我们在仿真环境和Unitree G1人形机器人上对物体拾取、搬运和放置任务进行了框架评估。该系统成功执行了单手和双手操作,并能够泛化到超出训练参考数据的放置位置。

英文摘要

Humanoid loco-manipulation requires robots to interpret task instructions and scene semantics while executing coordinated whole-body motions. We propose a hierarchical framework that uses motion keyframes as an intermediate representation between Vision-Language Model (VLM) planning and Reinforcement Learning (RL) control. Each keyframe specifies a target whole-body robot pose and, when applicable, an object pose. Given a language instruction, scene observations, and execution feedback, the VLM selects successive task-relevant keyframes from a predefined library. The selected keyframes are retargeted to the current scene to account for object poses and dimensions. A keyframe-conditioned whole-body policy then generates joint-level actions to reach these goals. We introduce a saliency-based keyframe sampling strategy for low-level policy training that improves end-to-end task success rate from 44% to 92% when using sparse VLM keyframes. We evaluate our framework on object pickup, transport, and placement tasks in simulation and on a Unitree G1 humanoid. The system successfully performs both one- and two-handed manipulation and generalises to placement locations beyond the training reference data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑