发表机构
JIUTIAN Research; Zhongguancun Academy(中移九天; 中关村学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出World Tokens架构,通过训练时的World Adapter衔接视觉-语言理解、世界动态建模与动作生成,在保持高效部署的同时提升具身策略性能,在多个基准测试中取得优异结果。
AI 中文摘要
视觉-语言-动作(VLA)模型是具身策略中广泛采用的范式,擅长高效的闭环控制,但未明确建模任务执行过程中物理场景的演变。新兴的世界-动作模型(WAMs)利用预训练的视频世界模型捕捉时空演变,不过在控制循环中保留未来生成或大型视频骨干网络会大幅增加推理成本。本文提出World Tokens,一种围绕World Adapter构建的具身策略架构,该适配器衔接视觉-语言理解、世界动态建模与动作生成,在训练阶段利用世界建模增强动作策略,同时保持高效部署。具体而言,World Adapter将视觉-语言模型(VLM)特征转换为固定数量的世界令牌,这些令牌为共同微调的未来视频去噪器提供条件,同时作为动作专家唯一的视觉-语言上下文。这种共享条件使未来视频去噪的梯度能够直接塑造动作预测所用的表示,而专属路由机制防止策略绕过该表示。部署时,移除世界模型分支,仅保留VLM、World Adapter和动作专家,无在线视频模型推理。采用2B骨干网络且未进行具身动作预训练的World Tokens,在LIBERO上表现极具竞争力,在SIMPLER上达到报告的最佳平均性能,相较匹配的仅动作基线大幅提升了现实世界R1 Pro的成功率,且每个动作块的生成延迟达到VLA级别。
英文摘要
Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes evolve as a task unfolds. Recently emerging world-action models (WAMs) leverage pretrained video world models to capture spatiotemporal evolution, yet retaining future generation or a large video backbone in the control loop substantially increases inference cost. We introduce World Tokens, an embodied policy architecture built around a World Adapter that bridges visual-language understanding, world-dynamics modeling, and action generation. It uses world modeling during training to enhance the action policy while preserving efficient deployment. Specifically, the World Adapter transforms VLM features into a fixed set of world tokens, which condition a jointly fine-tuned future-video denoiser and simultaneously serve as the action expert's sole visual-language context. This shared conditioning allows gradients from future-video denoising to directly shape the representation used for action prediction, while exclusive routing prevents the policy from bypassing that representation. At deployment, the world-model branch is removed, leaving only the VLM, World Adapter, and action expert, with no online video-model inference. With a 2B backbone and no embodied action pretraining, World Tokens is highly competitive on LIBERO, attains the best reported averages on SIMPLER, substantially improves real-world R1 Pro success over a matched action-only baseline, and generates each action chunk at VLA-level latency.