arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27753cs.CV

AWM-VLA:面向高效可解释视觉-语言-动作策略的对齐世界建模

AWM-VLA: AlignedWorld Modeling for Efficient and Explainable Vision-Language-Action Policies

An Lanji, Dawei Liu, Jin Li, Haoran Xu, Mei Chen, Yu Tian

首次发表
浏览论文内容

中文总结 AI 辅助

AWM-VLA在扩散-Transformer策略内嵌入对齐世界建模,通过未来令牌对齐与对象级语义预测,在RoboCasa等基准上成功率提升高达21%,实现高效可解释的机器人操作。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型已成为通用机器人操作的有力范式,然而它们往往是被动式的:策略直接将当前观测映射为动作块,而不对其决策的长期后果进行推理。先前赋予策略世界模型的尝试要么在像素空间中重建未来帧——成本高昂且被与任务无关的细节所主导——要么将世界模型与策略解耦,从而削弱了控制能力。我们提出AWM-VLA,一个将对齐世界建模直接嵌入扩散-Transformer策略内部的统一框架。遵循未来潜在表征对齐(FLARE)原则,我们添加可学习的未来令牌,其中间激活与未来观测的视觉-语言嵌入对齐,使策略在生成动作的同时能够预测长期后果。我们从两个方面扩展了这一范式。首先,我们引入一个以对象为中心的分离对齐目标,在全局未来嵌入之外预测未来的对象级语义,从而提高可解释性和多指令泛化能力。其次,我们通过原则性的加权方式平衡全局和对象中心对齐项与动作流匹配损失,从而在准确性与可解释性之间实现可控的权衡。在RoboCasa和人形桌面操作基准上,AWM-VLA在成功率上比先前的VLA和世界模型基线高出最多21%,提升了对新物体和新指令的泛化能力,并生成了以对象为中心的推理依据,在83%的案例中受到人类评估者的偏好。我们的方法仅为策略添加少量可学习令牌,并与任何扩散或流匹配策略兼容,使得对齐世界建模成为通用操作中廉价且广泛适用的组件。

英文摘要

Vision-language-action (VLA) models have become a powerful paradigm for generalist robotic manipulation, yet they are often reactive: the policy maps the current observation directly to an action chunk without reasoning about the long-term consequences of its decisions. Prior attempts to endow policies with world models either reconstruct future frames in pixel space---expensive and dominated by task-irrelevant detail---or decouple the world model from the policy, weakening control. We present AWM-VLA, a unified framework that embeds aligned world modeling directly inside a diffusion-transformer policy. Following the Future Latent REpresentation Alignment (FLARE) principle, we add learnable future tokens whose intermediate activations are aligned with vision-language embeddings of future observations, enabling the policy to anticipate long-term consequences while generating actions. We extend this paradigm in two ways. First, we introduce an object-centric decoupled alignment objective that predicts future object-level semantics alongside the global future embedding, improving both interpretability and multi-instruction generalization. Second, we balance the global and object-centric alignment terms against the action flow-matching loss through a principled weighting, yielding a controllable accuracy--interpretability trade-off. On RoboCasa and humanoid tabletop manipulation benchmarks, AWM-VLA outperforms prior VLA and world-model baselines by up to 21% in success rate, improves generalization to novel objects and instructions, and produces object-centric rationales that are preferred by human raters in 83 of cases. Our approach adds only a few learnable tokens to the policy and is compatible with any diffusion or flow-matching policy, making aligned world modeling an inexpensive, broadly applicable component of generalist manipulation.

发表机构

  • University of Electronic Science and Technology of China(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑