发表机构
Microsoft Research Asia; KAIST; The University of Tokyo(微软亚洲研究院; 韩国科学技术院; 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VLaRL利用VLA内部潜在表示作为迁移接口,在仿真中训练残差RL并部署到真实机器人,无需真实世界RL,在四个接触丰富任务和两个VLA骨干上均提升成功率。
AI 中文摘要
视觉-语言-动作(VLA)模型提供了广泛的、由指令调节的操作行为,但在接触丰富的交互过程中,其物理执行可能仍然不精确。残差强化学习(RL)可以在保持VLA冻结的同时纠正此类错误,但真实机器人上的RL成本高昂且对安全性要求苛刻。我们提出了VLA潜在条件强化学习(VLaRL),该方法使得针对冻结VLA的残差RL能够在仿真中训练,并在真实机器人上部署,而无需真实世界的RL或在线自适应。关键挑战在于克服仿真与真实之间的视觉差距来迁移学习到的残差策略。VLaRL不要求像素级别的视觉对应,而是利用VLA内部的视觉-语言潜在表示来调节残差控制,并将其作为仿真到真实迁移的接口,同时学习一个轻量级映射器,将仿真衍生的潜在特征转换为接近真实潜在分布。在四个接触丰富的操作任务和两个VLA骨干网络上,VLaRL在所有任务-骨干组合中均提升了真实世界的成功率,而受控消融实验则证明了潜在条件调节和潜在对齐对于迁移仿真训练的残差控制的重要性。
英文摘要
Vision-language-action (VLA) models provide broad, instruction-conditioned manipulation behaviors, but their physical execution can remain imprecise during contact-rich interaction. Residual reinforcement learning (RL) can correct such errors while keeping the VLA frozen, but real-robot RL is costly and safety-critical. We propose VLA Latent-Conditioned RL (VLaRL), which enables residual RL for frozen VLAs to be trained in simulation and deployed on real robots without real-world RL or online adaptation. The key challenge is transferring the learned residual policy despite the visual gap between simulation and reality. Rather than requiring pixel-level visual correspondence, VLaRL uses the VLA's internal vision-language latent representation to condition residual control and as the sim-to-real transfer interface, and learns a lightweight mapper that transforms simulation-derived latents toward the real latent distribution. Across four contact-rich manipulation tasks and two VLA backbones, VLaRL improves real-world success in all task-backbone combinations, while controlled ablations demonstrate the importance of both latent conditioning and latent alignment for transferring simulation-trained residual control.