DeicticVLA:在单一视觉-语言-动作模型(VLA)中统一基于语言和指示手势的指令模式
DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA
浏览论文内容
中文总结 AI 辅助
本研究提出DeicticVLA,将三种指令模式统一于单一VLA,经仿真与真实任务验证,其在未见场景和新类别下的表现优于仅用语言指令的模型,为相关设计提供指导。
中文摘要 AI 辅助
视觉-语言-动作模型(VLAs)允许用户用自然语言指定操作任务,但在同类或相似外观的物体中区分目标或放置目标需要详细表述,而VLAs可能无法可靠使用这些表述。我们提出DeicticVLA,该模型通过文本提示补全和指示手势 grounding,将语言指令(LI)、视觉-语言指令(VLI)和视觉指令(VI)规范为文本提示和指示掩码,使单个预训练VLA能够处理所有三种指令模式。在共享骨干、演示和匹配训练步数的条件下,我们在仿真中比较了两种RGB视觉提示方法、两种分离通道掩码提示方法和三种训练策略。在两阶段训练下,四种提示方法在分布内取得高成功率,但在未见布局中使用指示掩码的能力存在差异。跨方法的训练策略 ablation 显示,两阶段训练提升了这种使用能力,同时保留第二阶段LI数据可减轻遗忘,且不会降低VLI和VI性能。在三个真实世界任务中,一个策略支持所有模式,VLI和VI在未见表述、外观变化和新物体下的表现优于LI,对于未见类别,两者均达到100%成功率,而联合训练的LI仅为16.7%。这些结果证明了统一的三模式接口,并为DeicticVLA的设计提供指导。
英文摘要
Vision-Language-Action models (VLAs) allow users to specify manipulation tasks in natural language, but distinguishing a target or placement goal among objects of the same category or similar appearance requires detailed expressions that VLAs may not use reliably. We propose DeicticVLA, which canonicalizes Language Instruction (LI), Vision-Language Instruction (VLI), and Visual Instruction (VI) into a text prompt and deictic masks through text-prompt completion and deictic gesture grounding, enabling a single pretrained VLA to handle all three instruction modes. With a shared backbone, demonstrations, and matched training steps, we compare two RGB visual prompting methods, two separate-channel mask prompting methods, and three training strategies in simulation. Under two-stage training, the four prompting methods achieve high in-distribution success but differ in their ability to use deictic masks in unseen layouts. Across methods, training-strategy ablations show that two-stage training improves such use, while retaining second-stage LI data mitigates forgetting without reducing VLI and VI performance. In three real-world tasks, one policy supports all modes. VLI and VI outperform LI under unseen expressions, appearance changes, and novel objects. For unseen categories, both achieve 100% success, compared with 16.7% for jointly trained LI. These results demonstrate the unified three-mode interface and guide DeicticVLA design.
发表机构
- The University of Osaka(大阪大学)
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。