MergeOver:面向递归视觉Transformer的训练后令牌合并方法
MergeOver: Post-Training Token Merging for Recursive Vision Transformers
浏览论文内容
中文总结 AI 辅助
MergeOver是将令牌合并集成到递归权重共享Transformer的训练后方法,在ImageNet-1K上实现了GPU和树莓派5上的内存、延迟优化,为两类技术结合提供基线。
中文摘要 AI 辅助
视觉Transformer(ViTs)在计算机视觉领域展现出卓越性能,但存在参数量大、计算复杂度为二次方的问题,严重限制了其在资源受限的边缘硬件上的部署。尽管递归权重共享可减少参数数量,令牌合并能缓解计算和内存瓶颈,但在无需昂贵重训练的前提下将这两种范式相结合并非易事,该交叉领域仍未得到充分探索。我们提出MergeOver,一种将令牌合并(ToMe)集成到递归权重共享的切片递归Transformer(SReT)中的训练后方法。MergeOver通过未合并跟踪栈、约束安全的合并率调整以及跨空间排列的同步令牌质量跟踪,解决了该集成的空间和合并约束。我们进一步采用分阶段单次调度,在每个阶段的第一个块执行令牌缩减,并在后续所有递归迭代中保持固定序列长度。在ImageNet-1K上的基准测试显示,我们选定的配置使Top-1准确率降低1.47个百分点;在GPU上,批次大小为1和16时,峰值激活内存分别减少37.3%和38.4%,吞吐量在批次大小为1时降低21.7%,但在批次大小为16时提升21.7%;在树莓派5(ARM CPU)上,批次大小为1和16时延迟分别降低2.4%和17.6%。这些结果表明,MergeOver可在不进行重训练的情况下,恢复递归权重共享引入的部分吞吐量和内存开销,并为令牌合并与分层递归Transformer的结合提供了基线。
英文摘要
Vision Transformers (ViTs) demonstrate exceptional performance in computer vision but suffer from large parameter counts and quadratic computational complexity, severely limiting their deployment on resource-constrained edge hardware. While recursive weight-sharing reduces parameter counts and token merging mitigates computational and memory bottlenecks, integrating these two paradigms without costly retraining is non-trivial, leaving this intersection largely unexplored. We propose MergeOver, a post-training approach that integrates Token Merging (ToMe) into the recursively weight-shared Sliced Recursive Transformer (SReT). Through an Unmerge tracking stack, constraint-safe merge-rate adjustment, and synchronised token-mass tracking across spatial permutations, MergeOver resolves the spatial and merging constraints of this integration. We further employ a stage-wise single-shot schedule that performs token reduction at the first block of each stage and maintains a fixed sequence length throughout its subsequent recursive iterations. Benchmarked on ImageNet-1K, our selected configuration reduces top-1 accuracy by 1.47 percentage points. On the GPU, it reduces peak activation memory by 37.3% and 38.4% at batch sizes 1 and 16, while throughput decreases by 21.7% at batch size 1 but increases by 21.7% at batch size 16. On a Raspberry Pi 5 (ARM CPU), it reduces latency by 2.4% and 17.6% at batch sizes 1 and 16. These results show that MergeOver can recover a meaningful part of the throughput and memory cost that recursive weight-sharing introduces, without retraining, and provides a baseline for combining token merging with hierarchical recursive transformers.
发表机构
- University of Twente(特文特大学)
- Faculty of Engineering Technology, University of Twente(特文特大学工程技术学院)
- Computer Architecture for Embedded Systems, University of Twente(特文特大学嵌入式系统计算机架构实验室)
机构由 AI 辅助整理,请以论文原文为准。