发表机构
The Hong Kong University of Science and Technology (Guangzhou); The Hong Kong University of Science and Technology; Nanyang Technological University; Zhejiang University(香港科技大学(广州); 香港科技大学; 南洋理工大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SLIP-VLA框架,通过单步潜在想象实现高效未来建模,在仿真和真实操作任务中达到最先进性能,推理仅需12毫秒。
AI 中文摘要
视觉-语言-动作模型在机器人操作中日益有效,但大多数模型直接从当前观测预测动作,未显式建模未来场景演变。近期方法引入未来预测以改善动作生成,但密集未来建模通常需要昂贵的迭代去噪,而单步替代方案可能逊于多步对应方法。为调和高效未来建模与强动作性能,我们提出SLIP-VLA,一种策略学习框架,为VLA模型配备单步潜在想象以实现未来感知动作预测。SLIP-VLA通过单次去噪更新获得时间上密集的未来潜在表示,并通过将中间潜在特征与未来几何和语义特征对齐来提升这些表示的感知充分性。我们进一步通过动作条件潜在世界建模和逆动力学建模提升其控制充分性,显式耦合潜在转换与机器人动作。SLIP-VLA在多种仿真基准和真实世界操作任务中取得最先进性能,而其单步潜在想象仅需12毫秒。
英文摘要
Vision-Language-Action models are increasingly effective for robotic manipulation, yet most predict actions directly from current observations without explicitly modeling future scene evolution. Recent methods introduce future prediction to improve action generation, but dense future modeling often requires expensive iterative denoising, while one-step alternatives can underperform their multi-step counterparts. To reconcile efficient future modeling with strong action performance, we present SLIP-VLA, a policy learning framework that equips VLA models with a Single-Step Latent Imagination for future-aware action prediction. SLIP-VLA obtains temporally dense future latent representations with a single denoising update, and we improve the perceptual sufficiency of these representations by aligning intermediate latents with future geometric and semantic features. We further improve their control sufficiency through action-conditioned latent world modeling and inverse dynamics modeling, explicitly coupling latent transitions with robot actions. SLIP-VLA achieves state-of-the-art performance across diverse simulation benchmarks and real-world manipulation tasks, while its single-step latent imagination takes only 12 ms.
CommentsProject page and demonstration videos: https://haoxuanxu1024.github.io/SLIP_VLA/