arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

V-JEPA Policy:在预测性视觉潜在空间上构建有效的世界-动作模型

V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents

Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li

arXiv 2609.37250首次发表:更新:

发表机构

Tsinghua University; Shanghai Jiao Tong University; Fudan University; University of Science and Technology of China; Washington University in St. Louis; The Institute of Artificial Intelligence, China Telecom (TeleAI)(清华大学; 上海交通大学; 复旦大学; 中国科学技术大学; 圣路易斯华盛顿大学; 中国电信人工智能研究院(TeleAI))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出V-JEPA Policy框架,利用冻结的V-JEPA 2.1编码器的预测性视觉潜在空间构建世界-动作模型,无需完整预训练生成模型,在多个基准上达到竞争性能,并证明预测性潜在空间在分布偏移下更有效,且可迁移野外视频知识。

AI 中文摘要

世界-动作模型(WAMs)将未来视觉状态预测与动作生成相结合。通过适应大规模预训练的视频生成器或图像编辑模型,近期一系列WAMs继承了预测知识及其学习所在的模型。我们探究由大规模预测性预训练诱导的预测性视觉潜在空间,是否能在不继承完整预训练视觉生成模型的情况下,为有效的WAM学习提供充分基础。为回答此问题,我们提出V-JEPA Policy,一个在冻结的V-JEPA 2.1编码器潜在空间上构建WAM的简单框架。一个指令条件化的未来潜在预测器和一个流匹配动作专家在下游单一阶段中从零开始联合学习,预测器的未来信息上下文键-值状态用于条件化动作生成。V-JEPA Policy总参数为0.9B,其中0.6B可训练,在LIBERO、LIBERO-Plus和RoboCasa-GR1上取得了与代表性WAM和视觉-语言-动作基线相当的性能。在同一下游框架和训练预算下比较视觉基础,发现V-JEPA潜在空间比判别性、重建性和视频理解导向的替代方案更有效,尤其在分布偏移下。超越任务特定学习,在无动作标签的DROID视频-指令对上预训练预测器并适配为WAM,在下游控制和分布外泛化上带来显著提升。这些发现共同确立了预测性视觉潜在空间作为从任务特定演示中有效学习WAM以及迁移从更广泛野外视频中获取的未来建模知识的基础。我们的代码可在该https URL获取。

英文摘要

World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor's future-informed context key--value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video--instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at https://github.com/breez3young/VJEPA-Policy.

Comments19 pages, 5 figures, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑