arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

剖析视觉Transformer(ViT)的表征结构:一项严谨的架构研究

Dissecting Representation Structure in Vision Transformers: A Rigorous Architectural Study

Kim-Cuc Nguyen, Ngai-Man Cheung

arXiv 2610.11205首次发表:更新:

发表机构

Singapore University of Technology and Design; Temasek Laboratories(新加坡科技设计大学; 淡马锡实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过对ViT不同架构尺度的严谨分析,揭示其表征与泛化的关系,提出缓解特征崩溃的方案、可靠泛化预测指标等,所提代理指标提升了相关性排名,可指导高效ViT设计。

AI 中文摘要

表征结构对于理解视觉Transformer(Vision Transformer,ViT)架构及其泛化行为至关重要。然而,现有研究既未分离分析模块级特征,也未探究这些特征的交互如何影响性能估计。本研究首次对不同架构尺度下的特征信息开展严谨分析,实证揭示了ViT表征与泛化行为之间的关系,并利用这些见解指导高效ViT的设计。我们的贡献分为五个部分:在不同架构尺度下,1)我们识别出初始化时的特征崩溃现象,该现象会导致冗余,并提出一种缩减方案来缓解此问题;2)我们使用熵和最小特征值量化特征信息,证明这些指标可作为泛化预测的可靠指标;3)我们表明,token空间中的特征比嵌入空间中的特征提供更忠实的表征;4)我们发现一个意外结果:ViT层内线性子模块产生的特征对泛化性能的预测至关重要;5)我们提出的代理指标相较于现有基线,相关性排名提升了18-48%,且能有效识别在更低或相当计算成本下实现更高准确率的ViT架构。

英文摘要

Representation structure is crucial for understanding Vision Transformer (ViT) architectures and their generalization behavior. However, prior studies neither isolate nor analyze module-level features nor investigate how their interactions contribute to performance estimation. In this work, we conduct the first rigorous analysis of feature information across diverse architectural scales, empirically uncover the relationship between ViT representation and generalization behavior, and leverage these insights to guide efficient ViT design. Our contributions are fivefold: Across diverse architectural scales, 1) We identify feature collapse at initialization, which leads to redundancy, and propose a reduction scheme to mitigate this issue. 2) We quantify feature information using entropy and the minimum eigenvalue, demonstrating that these metrics serve as reliable indicators for generalization prediction. 3) We show that feature in the token space provides a more faithful representation than those in embedding space. 4) We discover an unexpected finding: features produced by linear submodules within ViT layers are critical for the prediction of generalization performance. 5) Our proposed proxy improves the correlation ranking by 18-48% over prior baselines and can effectively identify ViT architectures that achieve higher accuracy at lower or comparable computational cost.

CommentsAccepted in IEEE VCIP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑