arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

世界动作智能体:通过世界动作预演利用视觉语言模型进行机器人操作

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

Yehang Zhang, Haojian Huang, Yifan Chang, Jianchong Su, Bohan Zhou, Yingjie Xu, Wosong Chen, Tianhao Zhou, Chenxu Wang, Tianyi Zhang, Yangkai Wei, Wenqian Li, Shiyuan Deng, Yinchuan Li, Ying-Cong Chen, Zexi Li

arXiv 2609.29964首次发表:更新:

发表机构

HKUST(GZ); CUHK; Knowin AI(香港科技大学(广州); 香港中文大学; Knowin AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出世界动作智能体(WAA),通过视觉动作工作空间实现视觉语言模型直接操控机器人,利用接触视图、动作预演和视图内校正,在LIBERO-Pro上达到75.6%成功率并提升小模型域外性能。

AI 中文摘要

通用视觉语言模型(VLM)为机器人操作带来了广泛的知识和空间推理能力,然而现有系统要么间接使用它们来预测约束或编写程序,要么仅向它们提供场景视图而非一个可供行动的世界。我们提出了世界动作智能体(WAA),这是一个多智能体框架,通过它视觉语言模型能够使用基本工具操控机器人,并在一个视觉动作工作空间内做出每一个决策。该工作空间具有三个特性。接触视图根据场景几何自动选择,呈现当前交互周围的场景。动作预演将每个动作转化为可编辑的提议,智能体(单独或通过想象智能体)在执行前根据规划反馈进行预览和修订。视图内校正闭环连接了观察、预演和低级执行,使智能体能够在观察到残余偏移的视图中消除这些偏移。通过同一工作空间,WAA以两种方式获取具身程序性知识:它从专家视频和人类教学中在基于证据的审查下演化多模态技能,并通过技能智能体进行咨询;其交互轨迹训练较小的视觉语言模型来操控同一框架。在LIBERO-Pro上,仅从LIBERO-90演化技能的WAA达到了最先进的75.6%平均成功率,超越了端到端视觉语言动作模型、代码即策略智能体以及具有相同骨干网络的视觉框架基线;相同的技能在无需进一步学习的情况下在robosuite上依然有效。对Qwen3.5-9B在框架轨迹上进行微调,将其域外成功率从1.7%提升至43.3%。

英文摘要

General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. The workspace has three properties. Contact views, selected automatically from the scene geometry, present the scene around the current interaction. Action rehearsal turns each action into an editable proposal that the agent, alone or through an Imagination Agent, previews and revises against planning feedback before execution. In-view correction closes the loop between observation, rehearsal, and low-level execution, letting the agent remove residual offsets in the view where it observes them. Through the same workspace, WAA acquires embodied procedural knowledge in two ways: it evolves multimodal skills from expert videos and human teaching under evidence-based review and consults them through a Skill Agent, and its interaction traces train smaller VLMs to pilot the same harness. On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone; the same skills remain effective on robosuite without further learning. Fine-tuning Qwen3.5-9B on harness traces raises its out-of-domain success from 1.7% to 43.3%.

CommentsWorking in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑