arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38078cs.RO

MotorMind:为通用视觉语言模型搭建零样本机器人操作脚手架

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

MotorMind通过中层动作表示和异步执行框架,使通用VLM无需外部模型即可实现零样本机器人操作,在LIBERO-PRO上达到66.7%成功率,在真实xArm6上达到95%平均成功率。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型推动了机器人操作的发展,但它们在面对新任务和新环境时的零样本泛化能力仍然有限,而且它们对专门训练的依赖使其无法直接受益于快速发展的通用视觉语言模型(VLM)。与此同时,最近的智能体机器人系统利用VLM进行高层推理,或利用编码智能体进行机器人控制,但往往依赖大量外部模型和工具,引入了额外的复杂性和成本。这促使我们提出一个问题:通用VLM能否本身就像人类远程操作员一样,通过直接观察进行推理、发出动作并持续适应执行反馈来操作机器人,而无需依赖外部模型(如学习到的动作专家、编码智能体)或像SAM3这样的接地工具?在这项工作中,我们引入了MotorMind,一个机器人操作框架,它将VLM提出的中层动作连接到确定性机器人控制和反馈,并配备异步监控和后台记忆更新。在没有任务特定策略训练、编码智能体或像SAM3这样的额外接地工具的情况下,MotorMind在基础LIBERO-PRO套件上实现了66.7%的成功率,在扰动条件下实现了53.8%的成功率,而我们评估的先前零样本方法分别最多只有13.3%和19.2%。相同的接口在真实xArm6机器人上,在直接操作和人为扰动设置下平均成功率达到了95%。用更强的VLM替换骨干网络进一步提高了性能,而剩余的错误——主要源于视觉接地、具身推理和动作知识——随着VLM能力的提升而减少。这些结果表明,通用VLM在配备适当的中层动作表示和异步执行框架时,能够执行有效的零样本机器人操作。

英文摘要

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost. This motivates us to ask: Can a general-purpose VLM itself operate a robot more like the human teleoperator by reasoning directly from observations, issuing actions, and continuously adapting to execution feedback, without relying on external models such as learned action experts, coding agents or grounding tools like SAM3? In this work, we introduce MotorMind, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates. Without task-specific policy training, coding agents, or additional grounding tools such as SAM3, MotorMind achieves 66.7% success on the base LIBERO-PRO suites and 53.8% under perturbations, compared with at most 13.3% and 19.2%, respectively, for the prior zero-shot methods we evaluate. The same interface reaches 95% average success on a real xArm6 robot across direct manipulation and human-perturbation settings. Replacing the backbone with a stronger VLM further improves performance, while the remaining failures - primarily due to visual grounding, embodied reasoning, and action knowledge - decrease as VLM capability improves. These results show that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchronous execution harness, can perform effective zero-shot robotic manipulation.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

↑