arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29382cs.ROcs.LG

流匹配视觉-语言-动作模型中的解耦早期退出用于任务相关计算分配

Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs

Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro, Subramanian Ramamoorthy, Matteo Matteucci, Alessandro Suglia

首次发表
浏览论文内容

中文总结 AI 辅助

提出一种在流匹配VLA中联合配置骨干深度、动作专家深度和去噪步骤的框架,通过轻量退出Transformer和KV缓存合成实现任务相关计算分配,显著降低延迟和计算量并提升成功率。

中文摘要 AI 辅助

流匹配视觉-语言-动作(VLA)模型已成为通用机器人控制的潜在解决方案,其设计结合了预训练的视觉-语言模型(VLM)骨干与生成连续机器人动作的动作专家。尽管这些模型展现出令人印象深刻的能力,但由于其参数数量极高,其计算需求往往对机器人控制而言过于高昂。为缓解这些低效问题,现有方法主要通过在早期退出时跳过VLM骨干层或减少去噪步骤,而动作专家深度保持不变。我们提出一个框架,将骨干深度$V$、动作专家深度$A$和去噪步骤$D$作为VLA中三个可联合配置的计算轴。从预训练的VLA开始,我们在骨干和动作专家的中间深度附加轻量级退出Transformer(ET),训练其将策略的最后一层蒸馏到每个退出点。此外,我们引入一种KV缓存合成机制,管理被跳过的骨干层的缺失键和值,使动作专家能够比骨干更深地退出。最后,我们表明最优计算预算依赖于任务,不同任务受益于不同轴和深度。值得注意的是,我们的方法不需要从头训练原始策略,对于每个退出点,SmolVLA的参数仅增加$2.1\\%$,$\pi_{0.5}$增加$4.1\\%$。我们在两个流匹配VLA(SmolVLA,$\pi_{0.5}$)和两个基准(LIBERO,Meta-World)上验证了我们的方法,揭示了互补效应:$V$和$A$分别减少FLOPs和延迟,而$D$同时改善两者。我们的联合配置$(V,A,D)$将延迟降低$79.2\\%$,计算量(FLOPs)降低$31.8\\%$,同时平均成功率提高$5.6\\%$。

英文摘要

Flow-matching Vision-Language-Action (VLA) models have emerged as a potential solution for generalist robot control, designed by combining a pretrained Vision-Language Model (VLM) backbone with an action expert that generates continuous robot actions. While these models exhibit impressive capabilities, due to their very high number of parameters, their computational requirements are often prohibitive for robotics control. To mitigate these inefficiencies, existing methods predominantly skip VLM backbone layers with early exits or reduce denoising steps, while leaving action expert depth untouched. We propose a framework that exposes backbone depth $V$, action expert depth $A$, and denoising steps $D$ as three jointly configurable compute axes in a VLA. Starting from a pretrained VLA, we attach lightweight Exit Transformers (ET) at intermediate depths in both the backbone and the action expert, trained to distil the last layer of the policy into each exit. Furthermore, we introduce a KV Cache synthesis mechanism that manages the missing keys and values of the skipped backbone layers, allowing the action expert to exit deeper than the backbone. Finally, we show that the optimal compute budget is task-dependent, with different tasks benefiting from different axes and depths. Notably, our method does not require training the original policy from scratch, and for each exit, it increases the number of parameters by only $2.1\%$ for SmolVLA and $4.1\%$ for $π_{0.5}$. We validate our approach across two flow-matching VLAs (SmolVLA, $π_{0.5}$) and two benchmarks (LIBERO, Meta-World), revealing complementary effects: $V$ and $A$ respectively reduce FLOPs and latency, while $D$ improves both. Our joint configurations $(V,A,D)$ reduce latency by $79.2\%$ and computation (FLOPs) by $31.8\%$, while improving mean success rate by $5.6\%$.

发表机构

  • Politecnico di Milano(米兰理工大学)
  • University of Edinburgh(爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑