AI 中文总结
提出YUBI-STAG框架,自动丰富操作演示的交互语义,并通过蒸馏的YUBI-VLM实现高效视频标注,后训练VLA策略以增强细粒度语言指令遵循和接触感知能力。
AI 中文摘要
视觉语言动作(VLA)模型通过大规模预训练获得了广泛的操作能力,然而通过语言来激发这些能力需要指令与物理交互之间的细粒度对齐。现有的机器人演示通常仅提供粗略的任务描述,省略了动作的执行方式,包括哪个夹爪执行操作、接触了哪个物体,以及如何抓取和移动该物体。我们提出了YUBI-STAG,一个用于时空标注与接地(Spatio-Temporal Annotation and Grounding)的框架,该框架自动为操作演示丰富交互相关语义,以将预训练的VLA模型与细粒度操作语言对齐。YUBI-STAG结合接触物体分割与视觉语言模型,标注物体身份、属性和状态、每个夹爪的动作、双手协调以及空间接地的交互。为解决YUBI-STAG对局部化序列和多阶段VLM推理的依赖,我们将其蒸馏为YUBI-VLM。YUBI-VLM直接从原始未分割视频中通过少量推理调用恢复动作结构和标注,并仅依靠腕部视角运行。我们在YUBI-STAG-Bench上对两个框架进行了时间、语义和空间接地任务的评估。YUBI-VLM在减少推理调用和缩短运行时间的同时,保留了YUBI-STAG的大部分标注准确性,并能泛化到未见过的操作。最后,在这些标注上对VLA策略进行后训练,使其与细粒度语言和接触感知结构对齐。双手实验展示了改进的性能和指令遵循能力,包括对物体身份、执行夹爪、目标位置以及原始标签中缺失的空间关系的控制。
英文摘要
Vision-Language-Action (VLA) models acquire broad manipulation capabilities via large-scale pretraining, yet eliciting them through language requires fine-grained alignment between instructions and physical interactions. Existing robot demonstrations typically provide only coarse task descriptions, omitting how actions are executed, including which gripper acts, which object is contacted, and how it is grasped and moved. We introduce YUBI-STAG, a framework for Spatio-Temporal Annotation and Grounding that automatically enriches manipulation demonstrations with interaction-rich semantics to align pretrained VLAs with fine-grained manipulation language. Combining contact-object segmentation with vision-language models, YUBI-STAG annotates object identities, attributes and states, per-gripper actions, bimanual coordination, and spatially grounded interactions. To address YUBI-STAG's reliance on localized sequences and multi-stage VLM inference, we distill it into YUBI-VLM. YUBI-VLM directly recovers action structure and annotations from raw, unsegmented video in few inference calls and operates from wrist views alone. We evaluate both frameworks on YUBI-STAG-Bench across temporal, semantic, and spatial grounding tasks. YUBI-VLM retains much of YUBI-STAG's annotation accuracy with fewer inference calls and shorter runtime while generalizing to unseen manipulations. Finally, post-training VLA policies on these annotations aligns them with fine-grained language and contact-aware structure. Bimanual experiments demonstrate improved performance and instruction following, including control over object identity, acting gripper, target location, and spatial relations absent from original labels.
CommentsProject page: https://yubi-stag.airoa.io/