门控而非缓存:门控溯源界定无训练VLA令牌跳过的闭环可靠性
The Gate, Not the Cache: Gate Provenance Bounds the Closed-Loop Reliability of Training-Free VLA Token Skipping
浏览论文内容
中文总结 AI 辅助
本研究针对无训练VLA令牌跳过的闭环可靠性问题,提出动作松弛刷新方法,集成后修复门控自采集导致的性能崩溃,服务延迟降低18%-22%。
中文摘要 AI 辅助
令牌跳过是一种广泛应用的无训练加速视觉-语言-动作(VLA)模型的方法,在每个控制步骤中根据门控绕过大多数视觉令牌的计算。然而,当后续门控来自之前加速的前向传播时,某一步跳过的令牌也是后续门控最不关注的令牌,这种损伤会在多个控制步骤中累积,直至任务失败。本研究考察该类方法所基于的两种机制:复用(reuse)与删除(deletion),并在相同场景下对比门控信号来源对两种机制的影响。在LIBERO-Object数据集上,当跳过率为0.9时,若门控来自模型自身的加速前向传播,两种机制均会出现性能崩溃:复用机制下性能降至0.68,删除机制下降至0.31,而完整计算(dense)的性能为1.00,但所评估的动作级检测器无法察觉这种崩溃。区分崩溃与完整计算级运行的并非机制,而是门控是否干净——即由未跳过任何令牌的前向传播计算得到。因此,本研究提出动作松弛刷新(actuation-slack refresh):在机器人执行当前动作块的关键路径之外,运行一次完整前向传播,为下一步提供干净的门控和新鲜的KV缓存。由于所测检测器无法可靠揭示故障,该刷新为无条件执行而非触发式执行。两种机制的性能均恢复至0.98,同时保持了跳过的速度和完整前向传播的信息。随后,本研究将该刷新方法集成到两种VLA策略、4个LIBERO套件和4个SIMPLER任务的最先进缓存与剪枝方法中,成功修复了所有因使用自采集门控导致的崩溃。在仿真和物理机器人上的测量显示,服务延迟比完整计算降低了18%-22%。研究结论为:加速VLA的闭环可靠性由门控信号的来源决定,而非令牌跳过的方式。
英文摘要
Token skipping is a widely used training-free way to accelerate vision--language--action (VLA) models by bypassing computation for most visual tokens at each control step according to a gate. When the next gate is harvested from the previous accelerated forward, however, the tokens skipped at one step are also the ones least visible to the next gate, and the damage can compound across control steps until the task fails. We study the two mechanisms this class is built on, reuse and deletion, crossing each against where its gate signal comes from on identical episodes. At a skip ratio of 0.9 on LIBERO-Object, both collapse when the gate comes from the model's own accelerated forwards, to 0.68 under reuse and to 0.31 under deletion against a dense 1.00, and the collapse is invisible to the action-level detectors we evaluate. What separates collapse from dense-level operation is not the mechanism but whether the gate is clean, computed by a forward that skipped nothing. We therefore propose actuation-slack refresh, one dense pass run while the robot executes its current action chunk, off the critical path, that hands the next step a clean gate and a fresh KV base. Since the measured detectors do not reliably reveal the failure, the refresh is unconditional rather than triggered. Both mechanisms then recover to 0.98, keeping the speed of skipping and the information of a dense pass. We then integrate the refresh into state-of-the-art caching and pruning methods across two VLA policies, 4 LIBERO suites, and 4 SIMPLER tasks, where it repairs every collapse caused by using a self-harvested gate. Serve latency drops 18--22\% below dense, measured both in simulation and on a physical robot. Where the gate signal comes from, not how tokens are skipped, decides closed-loop reliability for accelerated VLAs.