基于动作缓存与优化的视觉-语言-动作模型免训练加速
ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement
- Institute of Science Tokyo(东京科学研究所)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对视觉-语言-动作模型动作头计算瓶颈问题,提出ActionCache即插即用外部缓存,利用紧凑多模态键存储中间动作,可跨不同场景检索重用,在模拟和真实环境实验中显著加速推理,提升任务成功率。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型是通用机器人操作的一种有前景的方法。基于流匹配的VLA模型因能生成精确平滑动作序列并捕捉多模态分布而成功,但动作头中的迭代去噪过程是主要计算瓶颈。为此提出ActionCache,一个即插即用的外部缓存,通过重用过去中间动作从目标动作附近热启动生成,减少推理延迟。实验表明ActionCache在低延迟下保持高任务成功率,分别使$\pi_{0.5}$和GR00T-N1.6加速达11.75倍和34.43倍。
英文摘要
Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow-matching-based VLA models have shown remarkable success due to their capability to generate precise and smooth action sequences and capture multimodal distributions. However, the iterative denoising process in the action head acts as a major computational bottleneck, posing a critical challenge for real-time deployment. To address this challenge, we propose ActionCache, a plug-and-play external cache that opportunistically reuses past intermediate actions to warm-start generations from the vicinity of target actions, drastically reducing the inference latency. Specifically, ActionCache stores the intermediate actions with compact multimodal keys, which enables retrieval from similar past contexts across different episodes or even different tasks. Experimental results in simulation and real-world environments demonstrate that ActionCache maintains high task success rates in a low-latency regime, achieving action head inference acceleration of up to $10.44\times$ and $40.17\times$ for representative flow-based VLA, $π_{0.5}$ and GR00T-N1.6, respectively.