arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CARE:为视觉-语言-动作推理的加速提供认证

CARE: Certifying Acceleration for Vision-Language-Action Inference

Rui Liu, Tong Zheng, Jindong Gu, Zhipeng Wang

arXiv 2610.08917首次发表:更新:

发表机构

University of Maryland, College Park; University of Oxford; Google(马里兰大学帕克分校; 牛津大学; 谷歌)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CARE通过配对轨迹提供有限样本保证,认证VLA推理加速器不超预算,在LIBERO上实现9-10.8倍加速并保留至少85.8%的参考成功回合。

AI 中文摘要

尽管视觉-语言-动作(VLA)模型发展迅速,但在每个控制步骤运行它们仍然代价高昂。先前的工作通过动作分块和视觉令牌剪枝等技术加速VLA推理,通常基于延迟和平均任务成功率进行评估。然而,加速可能会丢弃信息并破坏原始策略本可解决的任务,这一风险被平均指标所掩盖。由于动作偏差在闭环轨迹中会累积放大,任务失败仅在完整回合中可观察,因此衡量这些失败具有挑战性。为此,我们通过从相同初始条件出发的配对轨迹来定义加速引发的失败,跟踪参考策略成功而加速策略失败的情况。为应对这一问题,我们引入CARE,一种用于认证加速器选择的方法。CARE在校准集上使用配对轨迹,提供有限样本保证,确保加速引发的失败风险低于用户指定的预算。它部署最快的认证候选,若无候选合格则回退到参考策略。由于仅依赖最终结果和测量的计算量,CARE可不变地应用于各种加速机制,同时顺序测试和失败触发的参考轨迹使认证成本可控。在四个LIBERO套件上使用OpenVLA-OFT,CARE认证了9.0至10.8倍的加速,同时保证(在95%置信度下)至少85.8%的参考解决回合得以保留。在严格预算下,无保证的选择器在高达75%的试验中超出预算,而CARE保持在预算内,其顺序形式比穷举评估少使用78.9%的轨迹。CARE进一步推广到π0.5的流步减少,以及Crafter中的Qwen3.5-9B和Llama-3.1-8B智能体。

英文摘要

While vision-language-action (VLA) models have advanced rapidly, running them at every control step remains expensive. Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success. However, acceleration may discard information and break tasks the original policy would solve, a risk hidden by average metrics. Measuring these failures is challenging because action deviations compound over closed-loop trajectories, meaning task failure is only observable across full episodes. We therefore define an acceleration-induced failure via paired rollouts from identical initial conditions, tracking when the reference succeeds but the accelerated policy fails. To manage this, we introduce CARE, an approach for certified accelerator selection. CARE uses paired rollouts on a calibration set to provide finite-sample guarantees that acceleration-induced failure risk stays below a user-specified budget. It deploys the fastest certified candidate, falling back to the reference if none qualify. By relying only on terminal outcomes and measured compute, CARE applies unchanged across diverse acceleration mechanisms, while sequential testing and failure-triggered reference rollouts keep certification affordable. On four LIBERO suites with OpenVLA-OFT, CARE certifies $9.0$--$10.8\times$ speedups while guaranteeing (at $95\%$ confidence) that at least $85.8\%$ of reference-solved episodes are preserved. Under tight budgets, selectors without guarantees exceed the budget in up to $75\%$ of trials, whereas CARE stays within budget and its sequential form uses $78.9\%$ fewer rollouts than exhaustive evaluation. CARE further generalizes to flow-step reduction for $π_{0.5}$, and to Qwen3.5-9B and Llama-3.1-8B agents in Crafter.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑