何时剪枝与剪什么?面向高效VLA的阶段性视觉令牌剪枝
When and What to Prune? Stage-Aware Visual Token Pruning for Efficient VLA
浏览论文内容
中文总结 AI 辅助
提出SAPrune,一种无需训练的视觉令牌剪枝框架,通过校准集动态选择剪枝层并采用双路径规则,在VLA推理中剪除87.5%令牌,实现1.718倍加速且保持任务成功率。
中文摘要 AI 辅助
视觉令牌剪枝是加速视觉-语言模型的有效方式,尤其适用于视觉-语言-动作(VLA)推理,因为在预测机器人动作之前需要处理大量视觉令牌。现有的剪枝方法通常基于注意力分数或特征多样性来估计哪些令牌可以被剪除,保留那些被高度关注或与其他令牌视觉上不同的令牌。然而,大多数方法采用固定的剪枝计划,例如在预设层剪枝一次或在均匀间隔的层进行剪枝。这种计划对VLA模型可能存在风险,因为模型在早期层可能不知道哪些视觉区域对动作至关重要。最初看起来不重要的令牌可能在模型将视觉观察与语言指令结合后变得有用。在这项工作中,我们提出了SAPrune,一个无需训练的视觉令牌剪枝框架,用于高效的VLA推理。SAPrune不是固定层剪枝,而是使用一个小型校准集来观察动作到视觉注意力在各层间的变化,并在注意力模式变得更可靠后才选择剪枝层。在每个选定的层,SAPrune应用双路径剪枝规则:一条路径保护被高度关注的视觉令牌不被剪除,另一条路径防止有用的周围上下文被丢弃。在LIBERO、SIMPLER和真实机器人任务上的实验表明,SAPrune剪除了87.5%的视觉令牌,实现了高达1.718倍的推理加速,同时保持了具有竞争力的任务成功率。
英文摘要
Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually estimate which tokens can be pruned based on attention scores or feature diversity, retaining tokens that are either highly attended or visually different from others. However, most of them use fixed pruning schedules, such as pruning once at a preset layer or pruning at uniformly spaced layers. Such schedules can be risky for VLA models, because the model may not know which visual regions matter for the action in early layers. Tokens that look unimportant at first may become useful after the model combines visual observations with the language instruction. In this work, we propose SAPrune, a training-free visual token pruning framework for efficient VLA inference. Instead of pruning at fixed layers, SAPrune uses a small calibration set to observe how action-to-visual attention changes across layers, and chooses pruning layers only after the attention pattern becomes more reliable. At each selected layer, SAPrune applies a dual-path pruning rule: one path protects strongly attended visual tokens from pruning, while the other prevents useful surrounding context from being discarded. Experiments on LIBERO, SIMPLER, and real-world robotic tasks show that SAPrune prunes 87.5% of visual tokens and achieves up to 1.718x inference speedup while maintaining competitive task success rates.
发表机构
- University of Macau(澳门大学)
机构由 AI 辅助整理,请以论文原文为准。