将视觉语言模型的智能迁移到机器人控制
Transferring the Intelligence of VLMs to Robotic Control
浏览论文内容
中文总结 AI 辅助
本文提出RoboDawn接口,通过离散命令和上下文学习使VLM零样本或单样本控制机器人,在仿真和真实任务中超越专用策略,实现SOTA性能。
中文摘要 AI 辅助
人类可以无缝地适应物理世界和数字世界,这表明虽然在具身、环境和任务上存在数字到现实的差距,但人类智能本身可能跨越这一差距进行迁移。这自然引出一个基本问题:视觉语言模型(VLMs)的智能能否类似地从数字世界泛化到物理世界,用于机器人控制?我们通过RoboDawn来研究这个问题,RoboDawn是一种人类直观的接口,通过一组紧凑的离散平移、旋转和夹爪命令,将机器人控制暴露给一个智能体VLM。使用该接口,VLM以闭环方式控制机器人:它观察当前视觉状态,推理下一个动作,执行该动作,并根据结果状态调整后续决策。此外,我们引入了一种上下文学习(ICL)方案,该方案使用少量演示来使VLM在接口使用和任务解决策略两方面都得到基础指导。在RoboTwin 2.0 C2R和RoboDojo上的实验表明,RoboDawn在无需针对特定任务的机器人训练的情况下实现了强劲的性能。在零样本设置中,RoboDawn优于几种在基准特定机器人数据上训练的强策略,而单个上下文演示进一步带来了显著的性能提升,并建立了最先进(SOTA)的结果。在RoboTwin 2.0 C2R上,成功率从零样本的53.2%提升到单样本的73.6%,超过了坚实的基线π0.5(46.0%)。在RoboDojo上也观察到类似的提升,成功率从零样本的35.67%提升到单样本的47.17%。同一框架也迁移到真实世界机器人,在Franka上执行将块放入篮子和堆叠块的任务。
英文摘要
Humans can seamlessly adapt to both physical and digital worlds, suggesting that while a digital-to-real gap exists in embodiment, environment and task, human intelligence itself may transfer across this gap. This naturally raises a fundamental question: can the intelligence of vision-language models (VLMs) similarly generalize from the digital world to the physical world for robotic control? We investigate this question through RoboDawn, a human-intuitive interface that exposes robotic control to an agentic VLM through a compact set of discrete translation, rotation, and gripper commands. Using this interface, the VLM controls a robot in a closed loop: it observes the current visual state, reasons about the next action, executes it, and adapts subsequent decisions to the resulting state. Furthermore, we introduce an in-context learning (ICL) scheme that uses a few demonstrations to ground the VLM in both interface usage and task-solving strategies. Experiments on RoboTwin 2.0 C2R and RoboDojo demonstrate that RoboDawn achieves strong performance without task-specific robot training. In the zero-shot setting, RoboDawn outperforms several strong policies trained on benchmarkspecific robot data, while a single in-context demonstration further yields substantial performance gains and establishes state-of-the-art (SOTA) results. On RoboTwin 2.0 C2R, the success rate increases from 53.2% zero-shot to 73.6% one-shot, exceeding the solid baseline π0.5 (46.0%). Similar gains are observed on RoboDojo, where success rate improves from 35.67% zero-shot to 47.17% one-shot. The same framework also transfers to real-world robots, performing block-in-basket and block stacking on Franka.
发表机构
- Tsinghua University(清华大学)
- Tencent Hunyuan(腾讯混元)
机构由 AI 辅助整理,请以论文原文为准。