arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36808cs.ROcs.AI

Spotter:让具身模型主导,VLM 反思

Spotter: Let the Embodied Model Lead, and the VLM Reflect for It

发表机构格里菲斯大学 · 清华大学 · 腾讯
另 2 家 · 查看机构详情
  • Griffith University(格里菲斯大学)
  • Tsinghua University(清华大学)
  • Tencent(腾讯)
  • Fudan University(复旦大学)
  • Tongji University(同济大学)

机构由 AI 辅助整理,请以论文原文为准。

Long Li, Qichao Zhao, Yue Yang, Fan Xu, Zhe Wang, Alan Wee-Chung Liew, Chao Qu, Heng Tao Shen, Shirui Pan

首次发表
浏览论文内容

中文总结 AI 辅助

Spotter让具身模型主导执行,VLM并行监控并仅在出错时干预反思,显著提升机器人任务成功率并降低延迟。

中文摘要 AI 辅助

当前的具身模型不会对自身的失败做出反应,尽管刚刚出错的原因可以为下一次尝试提供小的调整,这种反思正是语言模型思维增益背后的机制。我们测试它们是否能修复一个已知错误,这需要产生修正并判断其是否正确。在失败处停止并允许重试时,它们很少能通过自身的随机性或对错误的语言描述来修复,而最佳N选择在失败后无法选出成功的候选。我们将此归因于仅在成功演示上训练以及输入过于狭窄无法显示出错原因,并得出结论:反思必须来自视觉语言模型(VLM),它接收更多信息,如情节历史和文本,且更具通用性。先前VLM主导的工作让VLM规划每一步并调用具身模型作为工具,将VLM置于关键路径上。我们提出Spotter,它反转了角色:具身模型主导并连续执行,而VLM并行运行,通过轻量级本地筛选器监控,仅在检测到错误时干预,反思并纠正错误,然后返回控制权。我们使用Qwen和GPT作为VLM运行Spotter,两者都改进了具身模型;使用GPT时,Spotter在RoboCasa上将Cosmos Policy和π0.5分别提高了5.6和7.5个百分点,在RoboTwin 2.0的Hard设置中将π0.5从47.2%提高到57.0%,在真实机器人上从53%提高到83%。由于VLM仅在确认错误时介入,使用Qwen的成功情节仅比单独使用具身模型多花13到16秒,比使用相同模型的VLM主导基线少约70%的时间。我们的代码可在https://这个URL获取。

英文摘要

Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the kind of reflection behind the gains of thinking in language models. We test whether they can repair a known error, which requires producing a correction and judging whether it is right. Stopped at a failure and allowed to retry, they seldom repair it through their own randomness or from a language description of the error, and best-of-N selection cannot pick the successful candidate after a failure. We attribute this to training only on successful demonstrations and to inputs too narrow to show what went wrong, and conclude that reflection must come from a vision-language model (VLM), which takes in far more information, such as the episode history and text, and is more general. Prior VLM-led work has the VLM plan every step and invoke the embodied model as a tool, placing the VLM on the critical path. We propose Spotter, which reverses the roles: the embodied model leads and executes continuously, while the VLM runs in parallel, monitors through a lightweight local screener, intervenes only when an error is detected, reflects on and corrects it, and returns control. We run Spotter with Qwen and with GPT as the VLM, and both improve the embodied models; with GPT, Spotter improves Cosmos Policy and $π_{0.5}$ by 5.6 and 7.5 percentage points on RoboCasa, and raises $π_{0.5}$ from 47.2% to 57.0% on the Hard setting of RoboTwin 2.0 and from 53% to 83% on a real robot. Because the VLM steps in only when an error is confirmed, a successful episode with Qwen takes only 13 to 16 s longer than with the embodied model alone and about 70% less time than with a VLM-led baseline using the same model. Our code is available at https://github.com/zqc3117/Spotter.

补充信息

↑