发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对分布偏移下视觉Transformer令牌缩减精度下降问题,提出后期集中的幂律缩减调度,在无额外计算下显著提升分布外精度,且跨方法、骨干和模态广泛有效。
AI 中文摘要
无训练令牌缩减通过跨层移除冗余令牌来加速视觉Transformer,以极小的计算代价恢复大部分原始精度。然而,这些方法主要针对干净数据进行设计和评估,在真实世界的分布偏移下,其与未压缩模型的精度差距随移除率增大而扩大。我们证明该差距由缩减调度(即移除的深度分布)决定,而该调度通常作为实现细节保持固定。具体而言,我们引入一个单参数、后期集中的幂律调度,在无额外推理成本的情况下持续提升分布外精度。在ImageNet-C上使用DeiT-S,后期调度在26%计算缩减下缩小了该差距的83%(+1.17个百分点),在较轻的7%缩减下缩小了99%(+0.26个百分点)。该增益不能归因于保留更多令牌或使用额外计算:在与平坦调度相同的计算量下,后期调度移除的令牌总数更多,最终保留的令牌更少,却仍然获胜。单层探针揭示了一种机制:早期缩减扰动会通过更多剩余层传播的特征,从而在深度上前置缩减误差。该效应广泛存在,适用于五种令牌缩减方法(ToMe、EViT、ATS、ATC、PiToMe)、九个骨干网络、所有ImageNet-C损坏类型、另外八个偏移套件,以及另外两种模态(视频和视觉-语言问答)。该效应还特定于偏移,在干净数据上仍为正向,并在最高严重级别5下单调上升至约4倍。该调度在六种测试时适应方法下保持其增益,且无需针对每个输入或每个域进行调优。
英文摘要
Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute. These methods, however, are designed and evaluated primarily on clean data, and under real-world distribution shift their accuracy gap to the uncompressed model widens with the removal rate. We show that this gap is governed by the reduction schedule, the depth profile of removal, usually left fixed as an implementation detail. Concretely, we introduce a one-parameter late-concentrated power-law schedule that consistently improves out-of-distribution accuracy over flat at no extra inference cost. On ImageNet-C with DeiT-S, the late schedule closes 83% of that gap at a 26% compute reduction (+1.17pp), and 99% of it at a lighter 7% reduction (+0.26pp). The gain cannot be attributed to retaining more tokens or using extra compute: held to flat's compute, the late schedule removes more tokens in total and leaves fewer tokens at the end, yet still wins. Single-layer probes point to a mechanism: earlier reductions perturb features that pass through more remaining layers, front-loading reduction error in depth. The effect is broad, holding across five token-reduction methods (ToMe, EViT, ATS, ATC, PiToMe), nine backbones, all ImageNet-C corruption types, eight further shift suites, and two further modalities, video and vision-language QA. It is also specific to shift, still positive on clean and rising monotonically to ~4x that at the highest severity 5. The schedule keeps its gain under six test-time adaptation methods, and needs no per-input or per-domain tuning.
Comments35 pages. Code: https://github.com/chahh9808/LaterIsBetter