IMLE-VLA:视觉-语言-动作策略的快速单步动作生成
IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies
- Simon Fraser University(西蒙弗雷泽大学)
- University of Pennsylvania(宾夕法尼亚大学)
- Alberta Machine Intelligence Institute(阿尔伯塔机器智能研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
IMLE-VLA通过条件隐式最大似然估计训练单步生成器替代迭代采样动作头,实现快速动作生成,在LIBERO基准上取得98.0%成功率并提升推理频率,同时保持鲁棒性。
AI中文摘要:
视觉-语言-动作(VLA)策略利用预训练的视觉-语言骨干网络实现强大的跨任务泛化能力。一种领先的设计将该骨干网络与通过扩散或流匹配训练的专用连续动作头相结合。然而,此类动作头依赖于迭代多步采样,例如在π_{0.5}中需要10步欧拉采样。这造成了推理瓶颈,导致机器人产生走走停停的运动并减慢任务完成速度。我们提出了IMLE-VLA,它将迭代动作头替换为通过条件隐式最大似然估计(cIMLE)训练的单步条件生成器。cIMLE目标促进了多模态动作覆盖,避免了朴素回归头的模式崩溃,同时完全消除了多步采样。当IMLE-VLA应用于π_{0.5}时,推理频率提高了3.67倍(55 Hz对比15 Hz),实现了高达11倍的动作吞吐量提升。在40个任务的LIBERO基准上,IMLE-VLA在所有基线中取得了最高的平均成功率(98.0%),同时推理频率领先。在LIBERO-plus的测试时扰动下,IMLE-VLA保持了π_{0.5}的鲁棒性,而其他基线性能急剧下降,证实了cIMLE动作头保留了泛化能力。在Franka Emika Panda上跨四个任务的真实世界实验表明,运动更平滑(抖动降低2.2倍至3.0倍),任务完成更快,IMLE-VLA在每个任务上都优于π_{0.5},并将每个回合的平均VLA推理时间减少了3.9倍至6.6倍。视频和代码可在以下网址获取:此https URL
英文摘要:
Vision-language-action (VLA) policies leverage pretrained vision-language backbones to achieve strong cross-task generalization. A leading design couples this backbone with a dedicated continuous action head trained via diffusion or flow matching. However, such heads rely on iterative multi-step sampling, for example 10 Euler steps in $π_{0.5}$. This creates an inference bottleneck that produces stop-and-go movement in the robot and slower task completion. We introduce IMLE-VLA, which replaces the iterative action head with a single-step conditional generator trained via conditional Implicit Maximum Likelihood Estimation (cIMLE). The cIMLE objective promotes multimodal action coverage, avoiding the mode collapse of naive regression heads while eliminating multi-step sampling entirely. When IMLE-VLA is applied to $π_{0.5}$, it increases inference frequency 3.67x (55 Hz vs. 15 Hz), enabling up to 11x higher action throughput. On the 40-task LIBERO benchmark, IMLE-VLA achieves the highest average success rate (98.0%) among all baselines while leading in inference frequency. Under the test-time perturbations of LIBERO-plus, IMLE-VLA retains $π_{0.5}$'s robustness while other baselines degrade sharply, confirming that the cIMLE head preserves generalization. Real-world experiments on a Franka Emika Panda across four tasks demonstrate smoother motion (2.2x to 3.0x lower jerk) and faster task completion, with IMLE-VLA outperforming $π_{0.5}$ on every task and reducing average VLA inference time per episode by 3.9x to 6.6x. Videos and code are available at https://kianhk6.github.io/IMLE-VLA/