arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于VLM智能体的上下文机器人学习

In-Context Robot Learning with VLM Agents

Dongzhou Cheng, Taoran Yi, Ye Fang, Xingwu Zhang, Fan Feng, Yixuan Li, Gengxiong Zhuang, Rongze Wang, Shuai Yang, Wei Song, Weizhi Xue, Minyan Wu, Jie Gui, Jiaqi Wang, Tong Wu

arXiv 2609.19138首次发表:更新:

AI 中文总结

针对机器人部署时泛化难题,提出GPT-Policy框架,利用VLM智能体结合上下文编译器与约束控制器实现无需梯度更新的上下文学习,实验证明其能提升真实机器人任务完成率。

AI 中文摘要

使机器人能够像人类一样轻松适应陌生环境,仍然是具身智能领域的一个登月目标。没有任何有限的演示集合能够覆盖机器人将遇到的所有任务和情况,这使得在部署时从上下文中学习的能力对于泛化至关重要。然而,这种上下文学习(ICL)在很大程度上仍然超出了现有机器人策略的能力范围。商业视觉语言模型(VLM)(如GPT-6 Astra)的广泛智能体能力提出了一个引人注目的问题:这些模型能否从演示、示例和交互反馈中学习,然后将这些信息转化为从新的初始状态出发的可执行且可验证的机器人行为,而无需梯度更新或对任务特定参数进行持久更改?我们引入了GPT-Policy,一个用于上下文机器人学习的通用智能体框架。GPT-Policy集成了一个上下文编译器,用于保留与任务相关的视觉转换;一个VLM,用于提出机器人工具动作;以及一个约束控制器,用于验证和执行每个动作并报告其结果。我们通过任务成功率和效率指标、跨模型的匹配比较以及受控的上下文消融来评估其可靠性和局限性。在真实机器人试验中,人类视频演示即使没有机器人动作标签也能提高任务完成率,而对齐的动作参考在接触敏感任务上带来了进一步的提升。这些发现将GPT-Policy定位为通过上下文学习实现机器人适应的一步,为将VLM的通用能力转化为物理行为提供了实证基础,并阐明了可靠部署必须克服的挑战。

英文摘要

Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment essential for generalization. Such in-context learning (ICL), however, remains largely beyond the reach of existing robotic policies. The broad agentic capabilities of commercial vision-language models (VLMs), such as GPT-6 Astra, raise a compelling question: can these models learn from demonstrations, examples, and interaction feedback, then translate that information into executable and verifiable robot behavior from a new initial state without gradient updates or persistent changes to task-specific parameters? We introduce GPT-Policy, a general-agent framework for in-context robot learning. GPT-Policy integrates a context compiler that preserves task-relevant visual transitions, a VLM that proposes robot-tool actions, and a constrained controller that verifies and executes each action and reports its outcome. We evaluate its reliability and limitations through task success and efficiency metrics, matched comparisons across models, and controlled context ablations. In real-robot trials, human video demonstrations improve task completion even without robot action labels, while aligned action references yield further gains on contact-sensitive tasks. These findings position GPT-Policy as a step toward robot adaptation through in-context learning, providing an empirical foundation for translating the general-purpose capabilities of VLMs into physical behavior and clarifying the challenges that must be overcome for reliable deployment.

CommentsProject Page: https://cheng-haha.github.io/GPT-Policy GitHub Code: https://github.com/cheng-haha/GPT-Policy

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑