arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PLaW-VLA:面向视觉-语言-动作策略的预测性潜在世界建模

PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies

Yu Liu, Hetian Guo, Tianlv Huang, Ziyi Cai, Wudi Chen, Hantang Wang, Qiutong Liu, Yingzhi Peng, Wei Han, Peijun Tang, Jianan Wang, Zipei Fan, Zhiyuan Zha, Xuan Song

arXiv 2610.12285首次发表:更新:

发表机构

Jilin University; Astribot; Harbin Institute of Technology, Shenzhen; The Hong Kong Polytechnic University; The University of Tokyo(吉林大学; Astribot; 哈尔滨工业大学(深圳); 香港理工大学; 东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PLaW-VLA是基于混合Transformer架构的VLA策略,通过在预训练预测表示空间建模任务相关未来状态,在RoboTwin和LIBERO-Plus数据集上分别实现长程控制和分布偏移泛化的性能提升,且推理延迟大幅降低。

AI 中文摘要

学习预测世界的演化可为视觉-语言-动作(VLA)策略提供长程控制所需的预测上下文,但其有效性取决于所建模的未来表示以及其如何调控动作生成。我们提出PLaW-VLA,该模型在预训练的面向预测的表示空间中建模与任务相关的未来状态,减少了对控制无关的视觉细节的预测需求。PLaW-VLA基于混合Transformer(Mixture-of-Transformers)架构构建,通过结构化因果注意力机制,将动作生成建立在观测历史、当前任务语义和预测未来状态的条件之上。实验结果显示,在RoboTwin Hard Horizon III上,相比反应式策略,其性能提升了11.8个百分点(pp);在零样本LIBERO-Plus上,相比面向重构的潜在预测,其性能提升了1.77个百分点,分别验证了其在长程控制和分布偏移下的泛化能力。通过避免低层次视觉重构,PLaW-VLA降低了未来预测的负担,实现了轻量级潜在世界模型,支持并行未来预测,在策略性能相当的情况下,其推理延迟仅为生成式世界-动作建模的约1/19。

英文摘要

Learning to predict how the world evolves can provide vision-language-action (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation. We introduce PLaW-VLA, which models task-relevant future states in a pretrained prediction-oriented representation space, reducing the need to predict control-irrelevant visual details. Built on a Mixture-of-Transformers architecture, PLaW-VLA conditions action generation on observation history, current task semantics, and predicted future states through structured causal attention. Experiments show a +11.8 percentage-point (pp) gain over reactive policies on RoboTwin Hard Horizon III and a +1.77 pp gain over reconstruction-oriented latent prediction on zero-shot LIBERO-Plus, supporting improved long-horizon control and generalization under distribution shift, respectively. By avoiding low-level visual reconstruction, PLaW-VLA lowers the burden of future prediction, enabling a lightweight latent world model with parallel future prediction and about 1/19 the inference latency of generative world-action modeling at comparable policy performance.

CommentsAccepted to the 10th Conference on Robot Learning (CoRL 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑