arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09940cs.ROcs.CV

Juno:驯化视觉-语言-动作模型的预测潜变量

Juno: Taming Predictive Latents for Vision-Language-Action Models

Yuchen Zhu, Chenyi Xu, Yulin Zhang, Gang Xu, Wentao Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

Juno提出一个基于动作条件JEPA的统一框架,通过控制对齐表示、解耦推理和测试时适应,解决了VLA模型预测潜变量的三大失败点,在SimEnv和真实机器人上显著提升成功率。

中文摘要 AI 辅助

联合嵌入预测架构(JEPAs)在表示空间中预测掩蔽或未来观测,为视觉-语言-动作(VLA)模型提供了天然的预测潜变量来源。然而,要使这些潜变量在预训练、策略学习和部署中发挥作用,需要解决三个失败点:与具身特定控制的不匹配、对动作学习的干扰、以及分布偏移下教师模型的校准错误。我们提出Juno,一个统一框架,围绕一个动作条件JEPA构建,该JEPA同时充当控制对齐的表示骨干、预测教师和可适应的动力学模型。在预训练期间,我们在具身匹配轨迹上训练它,并使用动态CLS损失将运动加权的补丁动力学转移到紧凑的全局状态。在策略学习期间,我们将当前帧的JEPA补丁融合到VLA感知中,并使用具有独立变换参数的解耦推理分支来蒸馏未来潜状态以生成动作。在部署期间,我们在所有观测到的转移(包括失败的 rollout)上适应世界模型,冻结适应后的教师,并使用LoRA适配器和可训练的动作头在验证的执行上重新对齐策略,无需专家修正或任务奖励。在SimplerEnv上,Juno将平均成功率从最强基线Qwen3GR00T的60.9%提高到68.5%,测试时适应进一步达到72.7%;在真实机器人上,在背景、高度和物体偏移下,它保持了70%–75%的成功率,而基础策略则崩溃至0%。

英文摘要

Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models. Yet making these latents useful across pretraining, policy learning, and deployment requires addressing three failures: mismatch with embodiment-specific control, interference with action learning, and teacher miscalibration under distribution shifts. We introduce Juno, a unified framework built around one action-conditioned JEPA that serves as a control-aligned representation backbone, a predictive teacher, and an adaptable dynamics model. During pretraining, we train it on embodiment-matched trajectories and use a dynamic CLS loss to transfer motion-weighted patch dynamics to a compact global state. During policy learning, we fuse current-frame JEPA patches into VLA perception and use a decoupled reasoning branch with separate transformation parameters to distill future latent states for action generation. During deployment, we adapt the world model on all observed transitions, including failed rollouts, freeze the adapted teacher, and re-align the policy on verified executions using LoRA adapters and a trainable action head, without expert corrections or task rewards. On SimplerEnv, Juno raises average success from $60.9\%$ to $68.5\%$ over Qwen3GR00T, the strongest baseline, and test-time adaptation further reaches $72.7\%$; on a real robot, it retains $70\%$--$75\%$ success under background, height, and object shifts where the base policy collapses to $0\%$.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Eastern Institute of Technology, Ningbo(宁波东方理工大学)
  • Hangzhou Dianzi University(杭州电子科技大学)
  • ShanghaiTech University(上海科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑