arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29208cs.ROcs.LG

AdaVLA:用于无需训练的视觉-语言-动作模型加速的自适应步长流匹配

AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models

Sunghwan Han, Youngtae Han, Youngmin Yi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出AdaVLA,一种无需训练的自适应框架,通过流匹配轨迹曲率度量动态调整推理步骤与MLP剪枝率,在LIBERO基准等测试中实现VLA模型加速且性能下降可忽略。

中文摘要 AI 辅助

基于视觉-语言模型(VLM)构建的视觉-语言-动作(VLA)模型,通过利用互联网规模的知识与多模态推理,显著提升了机器人的能力。然而,VLA模型密集的计算开销限制了其在设备上的部署,阻碍了对环境变化的实时响应。尽管已提出多种加速技术,但这些技术通常依赖微调或访问训练数据集,而由于隐私和专有性问题,训练数据集往往无法获取。此外,尽管基于流匹配的VLA模型已成为标准扩散模型的高效替代方案,但当前的加速工作主要针对VLM的推理成本,未能解决流匹配推理中固有的迭代常微分方程(ODE)求解过程。为解决这些局限,我们提出AdaVLA,一种用于快速且准确的基于流匹配的视觉-语言-动作模型的在线、无需训练的自适应框架。我们引入一种源自流匹配轨迹曲率的新型度量,用于在推理过程中量化动作生成置信度。该度量通过高效计算的重要性评估,实现推理步骤的动态减少和MLP剪枝率的自适应调整,无需访问训练数据。在LIBERO基准上使用Jetson AGX Orin设备进行的实验结果表明,我们的方法对π₀.₅和X-VLA分别实现了1.87倍和2.24倍的加速,且成功率下降可忽略不计。此外,我们使用SmolVLA在真实世界机器人任务上验证了我们方法的鲁棒性。

英文摘要

Vision-Language-Action (VLA) models, built upon Vision-Language Models (VLMs), have significantly enhanced robotic capabilities by leveraging internet-scale knowledge and multimodal reasoning. However, the intensive computational overhead of VLAs constrains on-device deployment, hindering real-time responses to environmental changes. While various acceleration techniques have been proposed, they often rely on fine-tuning or access to training datasets, which are frequently unavailable due to privacy and proprietary concerns. Moreover, although flow-matching-based VLAs have emerged as efficient alternatives to standard diffusion models, current acceleration efforts largely target VLM inference costs, failing to address the iterative ODE solving process inherent in flow matching inference. To address these limitations, we propose AdaVLA, an online, training-free adaptive framework for fast yet accurate flow-matching-based Vision-Language-Action models. We introduce a novel metric derived from the flow matching trajectory curvature to quantify action generation confidence during inference. This metric enables the dynamic reduction of inference steps and the adaptive adjustment of MLP pruning ratios through an efficiently computed importance evaluation, requiring no access to training data. Experimental results on the LIBERO benchmark using a Jetson AGX Orin device demonstrate that our method achieves $1.87\times$ and $2.24\times$ speedups for $π_{0.5}$ and X-VLA, respectively, with negligible degradation in success rates. Furthermore, we validate the robustness of our approach on real-world robotic tasks using SmolVLA.

发表机构

  • Sogang University(西江大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑