发表机构
The Hong Kong University of Science and Technology (Guangzhou); Ola Dimensions(香港科技大学(广州); 奥拉维度公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有世界-动作模型(WAM)存在的视频与指令语义不对齐问题,提出基于视觉-语言模型(VLM)的SG-WAM语义引导方法,通过注入文本接地且空间感知的语义前瞻,提升了WAM的指令遵循与操纵精度,经仿真和真实实验验证有效。
AI 中文摘要
世界-动作模型(World-Action Models,WAMs)已成为机器人操纵领域颇具前景的范式。然而,现有大多数WAM主要依赖视觉线索而非语言指令生成未来视频与动作,因为现成的文本编码器会独立于视觉观测嵌入指令,导致这些WAM预测的视频常与对应语言指令语义不对齐,降低了预测动作的准确性。为克服这一局限,我们提出SG-WAM,一种用于世界-动作模型的语义引导方法,利用视觉-语言模型(Vision-Language Model,VLM)作为语义规划器,增强世界-动作模型的指令接地能力。具体而言,我们训练一个基于VLM的规划器,以预测文本接地且空间感知的语义前瞻:文本接地的语义前瞻通过识别正确目标对象来接地指令,空间感知的语义前瞻则提供用于精确操纵的场景几何信息。随后,我们将该前瞻作为高层语义引导注入世界-动作模型,确保未来视频生成与动作预测均忠实地遵循语言指令。在仿真环境与真实世界中开展的大量实验表明,我们的语义引导方法具有优越性,展现出精确的操纵能力与强大的指令遵循能力。
英文摘要
World-Action Models (WAMs) have emerged as a promising paradigm for robotic manipulation. However, most existing WAMs generate future videos and actions by relying mainly on visual cues rather than language instructions, since off-the-shelf text encoders embed instructions independently of visual observations. As a result, the videos predicted by these WAMs are often semantically misaligned with their corresponding language instructions, which degrades the accuracy of the predicted actions. To overcome this limitation, we propose SG-WAM, a semantic guidance method for world-action models that leverages a vision-language model (VLM) as a semantic planner to enhance the instruction-grounding capacity of world-action models. Specifically, we train a VLM-based planner to predict text-grounded and spatial-aware semantic foresight. The text-grounded semantic foresight grounds the instruction by identifying the correct target objects, and the spatial-aware semantic foresight provides the scene geometry for precise manipulation. We then inject this foresight into the world-action model as high-level semantic guidance, ensuring that both future-video generation and action prediction faithfully follow the language instruction. Extensive experiments in simulation and the real world demonstrate the superiority of our semantic guidance method, showcasing precise manipulation and strong instruction-following capabilities.