Devol-ONE:一个自回归Transformer混合体,统一视觉-语言-动作与潜在世界建模
Devol-ONE: One Autoregressive Mixture of Transformers to Unify Vision-Language-Action and Latent World Modeling
浏览论文内容
中文总结 AI 辅助
Devol-ONE提出一种Transformer混合体架构,在单一自回归框架内统一视觉语言理解、潜在世界动态预测和动作生成,通过跨流联合预测和逐层注意力,在LIBERO等基准及真实机器人上验证了有效性。
中文摘要 AI 辅助
视觉语言动作(VLA)模型直接基于当前视觉和语言上下文来调节动作,而没有明确考虑场景在候选动作下如何演变。世界动作模型(WAM)试图通过预测未来状态来解决这一局限,但现有设计将预测和策略学习在架构上分开,仅通过预测输出(无论是像素空间的视频生成还是独立于策略训练的潜在预测模块)将它们连接起来。我们提出了Devol-ONE,一种Transformer混合体架构,在单一自回归框架内统一了视觉语言理解、潜在世界动态预测和动作生成。Devol-ONE不是将视觉语言标记编码一次并馈送给动作专家,而是跨视觉语言流和V-JEPA预训练的动态流联合运行自回归预测,在每一层关注视觉语言键值缓存,以在语言引导下预测未来潜在状态。动作专家因此持续受到语义推理和预测的物理动态的塑造,而不是由预先计算的固定表示所驱动。我们在LIBERO、LIBERO-PLUS、RoboTwin2.0上进行了广泛实验,并在Flexiv单臂和双臂设置上进行了真实世界评估。消融研究显示了动态流预测和逐层统一注意力的有效性,验证了我们模型的架构连贯性。
英文摘要
Vision Language Action (VLA) models condition actions directly on current visual and language context, without an explicit account of how the scene evolves under candidate actions. World Action Models (WAM) attempt to address this limitation by predicting future states, but existing designs keep prediction and policy learning architecturally separate, connecting them only through the predicted output, whether through pixel space video generation or a latent forecasting module trained independently of the policy. We present Devol-ONE, a Mixture of Transformers architecture that unifies vision language understanding, latent world dynamics prediction, and action generation within a single autoregressive framework. Instead of encoding vision language tokens once and feeding them to the action expert, Devol-ONE runs autoregressive prediction jointly across a vision language stream and a V-JEPA pretrained dynamics stream, attending to the vision language key-value cache at every layer to forecast future latent states under language guidance. The action expert is in turn shaped continuously by semantic reasoning and predicted physical dynamics rather than by a fixed representation computed in advance. Extensive experiments are conducted on LIBERO, LIBERO-PLUS, RoboTwin2.0 along with real-world evaluation on Flexiv single-arm and dual-arm setups. Ablation studies show the effectiveness of dynamic stream prediction and layer-wise unified attention to validate our model architectural coherency.