arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10522cs.ROcs.AIcs.CVcs.MM

Show-Harness:仅凭VLM智能体即可操控机器人

Show-Harness: Just a VLM Agent Can Play Robots

Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin, Yiqi Lin, Guoqiang Liang, Kevin Yuchen Ma, Qiming Huang, Mike Zheng Shou

首次发表
浏览论文内容

中文总结 AI 辅助

Show-Harness通过语义动作接口连接VLM意图与机器人动作,实现零样本控制及低成本微调,并扩展GUMI界面,实验证明其跨任务、具身和环境泛化优于现有范式。

中文摘要 AI 辅助

基础视觉语言模型(VLM)展现出关于世界的广泛智能,但将这种智能转化为机器人控制仍然具有挑战性。我们提出了Show-Harness,一种具身操控框架(Embodied Harness),通过连接意图与动作的紧凑语义接口,使VLM能够“操控”机器人。Show-Harness提供了离散的语义动作单元,VLM可以自然地对其进行推理,而具身特定的解释器则确定性地将这些动作单元落地为局部机器人动作,使VLM直接负责细粒度的物理决策。通过同一接口,Show-Harness展示了以下可行性:(1)直接解锁闭源前沿VLM用于零样本机器人控制;(2)仅需数GPU小时的微调,即可适配小型开源VLM用于低成本部署。我们进一步开发了GUMI(图形用户界面操作接口),将相同的语义动作空间扩展到基于GUI的演示收集,使人类和智能体无需专用遥操作硬件即可跨具身“操控”机器人。大量实验表明,配备Show-Harness的VLM智能体在任务、具身和环境间具有稳健的泛化能力,优于代表性的智能体范式和视觉-语言-动作(VLA)范式。这些结果表明,正确的接口可以在不增加额外模型容量或昂贵的具身特定预训练的情况下,释放基础VLM的大量具身能力。

英文摘要

Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness demonstrates the feasibility of (1) directly unlocking closed-source frontier VLMs for zero-shot robot control, and (2) adapting small-scale open-source VLMs for low-cost deployment with just a few GPU-hours of fine-tuning. We further develop GUMI (GUI Manipulation Interface), which extends the same semantic action space to GUI-based demonstration collection, allowing humans and agents to "play" robots across embodiments without specialized teleoperation hardware. Extensive experiments show that Show-Harness-equipped VLM agents generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms. These results suggest that the right interface can unlock substantial embodied capability from foundation VLMs, without requiring additional model capacity or costly embodiment-specific pretraining.

补充信息

↑