arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向视觉Transformer的解码器无关令牌合并:G2TM的系统性研究

Decoder-Agnostic Token Merging for Vision Transformers: A Systematic Study of G2TM

Victor Bercy, Martyna Poreba, Michal Szczepanski, Samia Bouchafa

arXiv 2609.18279首次发表:更新:

发表机构

Université Paris-Saclay; CEA; Univ Evry; IBISC(巴黎-萨克雷大学; 法国原子能委员会; 埃夫里大学; IBISC实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统评估了图引导令牌合并(G2TM)在多种分割框架、解码器及图像分类中的表现,证明其有效性源于编码器而非解码器,并发现其最优超参数主要取决于骨干网络预训练方案和目标数据集。

AI 中文摘要

视觉Transformer(ViTs)凭借自注意力机制,在一系列计算机视觉任务中取得了最先进的性能。然而,其复杂度随令牌数量呈二次方增长,这仍是ViT效率和规模化部署的主要障碍。令牌合并通过聚合冗余令牌来降低这一成本。然而,现有方法通常仅在单一架构中进行评估,因此尚不清楚其有效性是源于合并机制本身,还是源于与之搭配的特定解码器。我们将图引导令牌合并(G2TM)——一种插入ViT网络前端的单一模块——扩展到其最初的Segmenter设置之外。我们在三个语义分割框架(Segmenter、SETR、EoMT)和三个解码器家族(基于线性、基于Transformer、基于卷积)以及标准ViT图像分类中评估了G2TM。我们的结果表明,对于给定的骨干网络规模,G2TM的行为和精度-效率权衡在所有测试架构中保持一致,这表明其有效性是编码器的属性,而非解码器的属性。G2TM也能很好地泛化到图像分类,与语义分割相比,其精度下降幅度更小。我们进一步发现,G2TM的最优超参数主要取决于骨干网络的预训练方案和目标数据集,而非解码器的选择,这些超参数在ADE20K数据集上的分割模型中实现了GFLOPs一致降低22-47%,吞吐量最高提升74%。

英文摘要

Vision Transformers (ViTs) have achieved state-of-the-art performance across a range of computer vision tasks, mainly thanks to the self-attention mechanism. However, its complexity, increasing quadratically with the number of tokens, remains the major obstacle to ViT efficiency and deployment at scale. Token merging reduces this cost by aggregating redundant tokens. Yet existing methods are typically evaluated within a single architecture, leaving open whether their effectiveness stems from the merging mechanism itself or from the specific decoder they are paired with. We extend Graph-Guided Token Merging (G2TM), a single module inserted early in a ViT-based network, beyond its original Segmenter setting. We evaluate G2TM across three semantic segmentation frameworks (Segmenter, SETR, EoMT) and three decoder families (Linear, Transformer-, convolution-based), as well as standard ViT image classification. Our results show that G2TM's behavior and accuracy-efficiency trade-off are consistent across every tested architecture for a given backbone size, indicating that its effectiveness is a property of the encoder rather than the decoder. G2TM also generalizes well to image classification, achieving an even smaller degradation in accuracy compared to semantic segmentation. We further find that G2TM's optimal hyperparameters, resulting in a consistent drop in GFLOPs of 22-47% and an increase in throughput by up to 74% for segmentation models on ADE20K dataset, depend primarily on the backbone's pre-training recipe and on the target dataset, rather than on the decoder choice.

CommentsExtended version of https://cea.hal.science/cea-05578363, to be published in Communications in Computer and Information Science (CCIS), Springer. Codes are available at https://github.com/vbercy/g2tm

DOI:10.5220/0014267600004084

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑