DORA:用于视觉Transformer令牌剪枝的动态在线强化智能体
DORA: Dynamic Online Reinforcement Agent for Token Pruning in Vision Transformers
浏览论文内容
中文总结 AI 辅助
针对视觉Transformer令牌冗余问题,提出DORA,一个学习输入自适应剪枝策略的在线强化智能体,通过分层决策和影子评估,在保持精度的同时显著降低计算量。
中文摘要 AI 辅助
视觉Transformer(ViTs)在令牌数量上产生二次方自注意力成本。大多数令牌缩减方法在预设的逐层压缩调度中调整令牌身份,或离线搜索静态掩码,从而限制了在线调整剪枝时机和剪枝量的能力。我们提出DORA(动态在线强化智能体),它针对冻结的ViTs学习输入自适应的剪枝策略。在每个符合条件的块中,一个分层行动者决定是否剪枝、移除多少令牌以及从每张图像不断演化的表示中移除哪些令牌。由于早期删除会改变后续决策所观察到的状态,DORA将剪枝建模为有限时域马尔可夫决策过程。完整前缀影子评估将最终预测保真度转化为局部逐步奖励,而闭环精度反馈将保真度惩罚调整至共享的精度下降目标。特权评论家及所有影子计算仅用于训练。部署时保留冻结的主干网络和一个轻量级行动者,该行动者应用硬删除和打包的可变长度FlashAttention,将令牌缩减转化为实测加速。在ImageNet-1K上使用DeiT-Base,DORA相对于未压缩主干网络减少了38.4%的FLOPs,精度损失在一个百分点以内。在匹配精度下,跨四种ViT类型主干网络平均,DORA比相应的每主干网络基线均值少使用13.2%的FLOPs,吞吐量提高32.4%。在零样本迁移到ImageNet-A时,这些增益分别扩大到20.3%和45.6%。
英文摘要
Vision Transformers (ViTs) incur quadratic self-attention cost in the number of tokens. Most token-reduction methods adapt token identities within a prescribed layer-wise compression schedule, or search a static mask offline, and thus limit online adaptation of when and how much to prune. We propose DORA (Dynamic Online Reinforcement Agent), which learns an input-adaptive pruning policy itself for frozen ViTs. At each eligible block, a hierarchical actor decides whether to prune, how many tokens to remove, and which tokens to remove from each image's evolving representation. Because early deletions change the states observed by later decisions, DORA formulates pruning as a finite-horizon Markov decision process. Complete-prefix shadow evaluations convert final-prediction fidelity into localized per-step credit, while closed-loop accuracy feedback adjusts the fidelity penalty toward a shared accuracy-drop target. A privileged critic and all shadow computations are training-only. Deployment retains the frozen backbone and a lightweight actor that applies hard deletion and packed variable-length FlashAttention, converting token reduction into measured speedups. On ImageNet-1K with DeiT-Base, DORA reduces FLOPs by 38.4% relative to the uncompressed backbone within one percentage point of accuracy loss. Averaged across four ViT-type backbones at matched accuracy, DORA uses 13.2% fewer FLOPs and achieves 32.4% higher throughput than the corresponding per-backbone baseline means. Under zero-shot transfer to ImageNet-A, these gains widen to 20.3% and 45.6%, respectively.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。