arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人形机器人视觉语言动作闭环:用于可验证移动操作的持久3D对象令牌

Closing the Loop in Humanoid VLA: Persistent 3D Object Tokens for Verifiable Loco-Manipulation

Peng Ren, Haoyang Ge, Jiang Zhao, Cong Huang, Yukun Shi, Pei Chi, Kai Chen

arXiv 2607.18016首次发表:更新:

发表机构

BUAA; DeepCybo; ZGCI(北京航空航天大学; 深灵机器人; 未提及具体中文译名)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究人形机器人视觉语言动作中对象状态差异问题,提出持久对象令牌化方法(POT)并实例化为POT-VLA,通过RGB-D观察维护3D对象记录,转换为对象令牌用于动作生成与验证,实验表明该方法提升了任务成功率。

AI 中文摘要

视觉语言动作策略是通用机器人控制的一个有前景的基础,但长期的人形机器人移动操作要求机器人在运动、接触、遮挡和恢复过程中将任务对象视为持久的物理实体。我们将此问题研究为对象状态差异:用于调节全身动作的对象状态可能与用于判断动作是否实现预期物理关系的状态不同。我们提出了“持久对象令牌化”(POT),它从RGB-D观察中维护按角色索引的3D对象记录,并将其转换为用于全身动作专家的对象令牌。实例化为“POT-VLA”时,相同的对象记录可调节动作生成并支持几何谓词检查,产生一个闭环执行系统,其中对象状态既具有可操作性又具有可验证性。在Unitree G1上,POT-VLA将匹配的直接GR00T-N1.7基线在八个现实世界任务族中的成功率从39/80提高到71/80。在外部与Being-0对齐的参考中,POT-VLA在对齐的服务任务上实现了44/50的成功率,而Being-0论文报告的成功率为37/50。最大的收益出现在需要维持3D关系的任务上,这表明持久的以对象为中心的状态是可验证人形机器人视觉语言动作执行的有用抽象。

英文摘要

Vision-language-action policies are a promising foundation for general robot control, but long-horizon humanoid loco-manipulation requires the robot to treat task objects as persistent physical entities across movement, contact, occlusion, and recovery. We study this problem as object-state divergence: the object state used to condition a whole-body action can differ from the state used to decide whether the action achieved the intended physical relation. We propose \emph{Persistent Object Tokenization} (POT), which maintains role-indexed 3D object records from RGB-D observations and converts them into object tokens for a whole-body action expert. Instantiated as \emph{POT-VLA}, the same object records condition action generation and support geometric predicate checks, yielding a closed-loop execution system in which object state is both actionable and verifiable. On a Unitree G1, POT-VLA improves a matched direct GR00T-N1.7 baseline from 39/80 to 71/80 successes over eight real-world task families. In an external Being-0-aligned reference, POT-VLA achieves 44/50 successes on aligned service tasks, compared with the 37/50 success reported by the Being-0 paper. The largest gains occur on tasks requiring maintained 3D relations, suggesting that persistent object-centered state is a useful abstraction for verifiable humanoid VLA execution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑