AR-WAM:一种面向机器人操作的视觉条件智能体就绪世界动作模型
AR-WAM: A Visual-Conditioned Agent-Ready World Action Model for Robotic Manipulation
- The Hong Kong University of Science and Technology(香港科技大学)
- Shenzhen University of Advanced Technology(深圳理工大学)
- Sangfor Technologies Inc.(深信服科技股份有限公司)
- Mininglamp Technology(明略科技)
- The Chinese University of Hong Kong(香港中文大学)
- The University of British Columbia(不列颠哥伦比亚大学)
- Astribot
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
AR-WAM提出用视觉接地提示和可学习操作令牌替代语言指令,以0.5B参数模型在标准操作中匹配最强基线,并在长时任务上显著提升成功率。
AI中文摘要:
随着人工智能智能体能力的不断增强,智能体驱动的机器人控制正成为一种引人注目的范式。然而,现有的视觉-语言-动作(VLA)模型和世界动作模型(WAMs)仍然依赖自然语言指令来指定操作任务,这对于智能体驱动的控制来说是一种不合适的接口:指代模糊、空间不精确、与智能体固有的语言理解冗余,并且将意图与执行纠缠在一起。我们提出了AR-WAM,一种视觉条件化的、智能体就绪的世界动作模型,它用两种互补的条件取代了语言:一个视觉接地提示(目标物体的边界框)用于指示交互对象和位置,以及一个可学习的操作令牌用于指定要执行的原子技能。我们紧凑的0.5B参数模型,配有冻结的预训练视觉编码器且无语言编码器,在解码动作的同时在紧凑的潜在状态内预测场景演变,通过显式的、可监督的推理信号暴露策略的意图。一个模型无关的兼容层提供了三种原语(检测、执行和查询),使得本地VLM或在线智能体API可以直接驱动策略,而长时记忆和闭环错误恢复则委托给智能体侧。在RoboTwin 2.0、RMBench和真实的Astribot S1双臂平台上,AR-WAM在标准操作任务上匹配最强基线(平均成功率87.2%),并在依赖记忆和真实机器人长时任务上超越它们,成功率分别提高了5.9%和36.7%,同时保持了最低的推理延迟(14.1毫秒)。
英文摘要:
As AI agents become increasingly capable, agent-driven robotic control is emerging as a compelling paradigm. However, prevailing vision-language-action (VLA) models and world action models (WAMs) still rely on natural-language instructions to specify manipulation tasks, an ill-suited interface for agent-driven control: referentially ambiguous, spatially imprecise, redundant with the agent's inherent language understanding, and entangling intent with execution. We present AR-WAM, a visual-conditioned, agent-ready world action model that replaces language with two complementary conditions: a visual grounding prompt (a bounding box of the target) denoting the interaction object and location, and a learnable operation token dictating the atomic skill to execute. Our compact 0.5B-parameter model, with a frozen pretrained visual encoder and no language encoder, predicts scene evolution within compact latent states while decoding actions, exposing the policy's intent through explicit, supervisable reasoning signals. A model-agnostic compatibility layer provides three primitives (detect, execute, and query) so that local VLMs or online agent APIs can drive the policy directly, with long-horizon memory and closed-loop error recovery delegated to the agent side. On RoboTwin 2.0, RMBench, and a real Astribot S1 dual-arm platform, AR-WAM attains the highest average success on standard manipulation (85.7% over the clean and randomized settings) and outperforms all baselines on memory-dependent and real-robot long-horizon tasks, improving success rates by 5.9% and 36.7%, respectively, while maintaining the lowest inference latency (14.1 ms). Project page is at https://ar-wam.github.io/.