JEPA-WAM:通过JEPA潜在表示将生成的视觉指令连接到世界动作模型
JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations
AI总结:
JEPA-WAM通过为文本指令生成多样视觉指令,利用V-JEPA编码器提取任务级语义,提升世界动作模型的指令跟随能力,在真实机器人基准上显著优于现有方法。
AI中文摘要:
世界动作模型(WAMs)通过将动作专家与预训练的视频生成模型相结合,展示了强大的机器人操作能力。然而,当前的世界动作模型在仅依赖文本指令进行条件生成时,其指令跟随能力仍然有限。我们认为,这一局限性部分源于机器人学习数据中的结构性失衡:丰富的视觉-动作轨迹往往与稀疏且重复的语言标注配对,使得策略能够从视觉上下文和运动规律中识别任务,而非真正理解指令本身。为解决这一局限性,我们引入了JEPA-WAM,该方法为每条文本指令增加一组随机生成的视觉指令,为指令跟随提供多样化的视觉线索。具体而言,JEPA-WAM使用现成的文本到图像生成器,在无需训练生成器的情况下,基于文本指令采样多个任务完成图像。尽管这些生成的图像可能在外观和布局上与当前视觉场景不同,但它们与指令在语义上保持一致,并作为视觉目标参考。为了聚焦于外观之外的任务级语义,我们使用冻结的V-JEPA 2.1编码器对这些参考进行编码。得到的密集目标表示被压缩为紧凑的目标令牌,通过交叉注意力同时条件化视频和动作专家。我们进一步构建了一个真实机器人指令跟随基准,涵盖分布内、分布外场景和分布外指令设置。在该基准上,JEPA-WAM在这三种设置下的成功率分别达到87.3%、74.5%和80.9%,分别比π0和Fast-WAM高出至少10.0、27.3和14.5个百分点。
英文摘要:
World Action Models (WAMs) have demonstrated strong robotic manipulation capabilities by augmenting pretrained video generative models with action experts. However, current WAMs still show limited instruction-following ability when conditioned solely on text instructions. We argue that this limitation stems in part from a structural imbalance in robot-learning data: rich visual-action trajectories are often paired with sparse and repetitive language annotations, allowing policies to identify tasks from visual context and motion regularities rather than grounding the instruction itself. To address this limitation, we introduce JEPA-WAM, which augments each text instruction with a bank of stochastically generated visual instructions, providing diverse visual cues for instruction following. Specifically, JEPA-WAM uses an off-the-shelf text-to-image generator to sample multiple task-completion images conditioned on the text instruction, without training the generator. Although these generated images may differ from the current visual scene in appearance and layout, they remain semantically aligned with the instruction and serve as visual goal references. To focus on task-level semantics beyond appearance, we encode these references with a frozen V-JEPA 2.1 encoder. The resulting dense goal representations are compressed into compact goal tokens that condition both the video and action experts through cross-attention. We further construct a real-robot instruction-following benchmark covering in-distribution, out-of-distribution scene, and out-of-distribution instruction settings. On this benchmark, JEPA-WAM achieves success rates of 87.3%, 74.5%, and 80.9% in these three settings, outperforming π0 and Fast-WAM by at least 10.0, 27.3, and 14.5 percentage points, respectively.