发表机构
CASIA; UCAS; Alibaba Group; FiveAges(中国科学院自动化研究所; 中国科学院大学; 阿里巴巴集团; 五时代)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ViDAL通过动作变分自编码器将连续动作潜变量锚定于未来视觉动力学,构建视觉动力学接地动作潜空间,兼容多种VLA架构,在LIBERO、RoboTwin及真实机器人上显著提升成功率。
AI 中文摘要
视觉-语言-动作(VLA)模型已成为机器人策略学习的核心范式,其以三种形式预测动作:原始动作块、离散动作标记或连续动作潜变量。然而,现有的动作表示主要对动作轨迹进行建模,对由这些动作引起的视觉动力学考虑有限。我们提出ViDAL,一种视觉动力学接地的动作潜空间,将连续动作潜变量锚定在场景的未来视觉动力学中。具体而言,ViDAL通过训练一个动作变分自编码器(Action VAE)来重建动作块,同时将其潜变量与未来场景动力学对齐,从而学习动作潜空间。当集成到下游机器人策略中时,所提出的Action VAE作为即插即用的动作接口,兼容多种VLA架构,并支持可选的未来视频预测作为附加能力。实验上,ViDAL在LIBERO上以98.1%的平均成功率优于竞争基线,将RoboTwin 2.0上的多任务π0.5策略在50个双臂任务中的成功率从54.3%提升至65.5%(干净环境)和从33.2%提升至43.1%(随机环境),并在真实世界的单臂Franka和双臂ARX机器人平台上分别获得20.0%和23.4%的绝对成功率提升。
英文摘要
Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents. However, existing action representations primarily model action trajectories, with limited consideration of the visual dynamics induced by these actions. We introduce ViDAL, a Visual Dynamics-grounded Action Latent Space that anchors continuous action latents in the future visual dynamics of the scene. Specifically, ViDAL learns action latent space by training an Action Variational Autoencoder (Action VAE) to reconstruct action chunks while aligning its latent with future scene dynamics. When integrated into downstream robot policies, the proposed Action VAE serves as a plug-in action interface compatible with multiple VLA architectures and enables optional future-video prediction as an additional capability. Empirically, ViDAL outperforms competitive baselines on LIBERO with 98.1% average success, improves a multi-task $π_{0.5}$ policy on RoboTwin 2.0 from 54.3% to 65.5% (Clean) and from 33.2% to 43.1% (Random) success rates over 50 dual-arm tasks, and yields 20.0% and 23.4% absolute success-rate gains on real-world single-arm Franka and dual-arm ARX robot platforms.