Transformer几何观测站TGO-II:表征相似性观测站
Transformer Geometry Observatory TGO-II: Representational Similarity Observatory
- Department of Electronics Engineering Sardar Vallabhai National Institute of Technology (SVNIT), Surat, India(电子工程系萨达尔·瓦拉布希国家理工学院(SVNIT),印度萨拉特)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出TGO-II框架,通过CKA、SVCCA、TwoNN-ID和token协方差分析,发现ViT训练中表征专业化增强、内在维度增加且token交互结构保持,挑战了表征复杂性源于token独立的假设。
AI中文摘要:
尽管Vision Transformers在计算机视觉和语言应用中取得了显著成功,但其内部表征在训练过程中的几何演化仍未被充分理解。现有分析主要关注注意力机制和下游性能,而表征几何的演化在很大程度上未被探索。在这项工作中,我们提出了Transformer几何观测站-II(TGO-II),一个表征几何分析框架,旨在研究Transformer表征在监督训练过程中如何演化。TGO-II使用中心核对齐(CKA)、奇异向量典型相关分析(SVCCA)、二近邻内在维度(TwoNN-ID)和token协方差分析来分析Vision Transformer(ViT-Small/16)的表征。我们的实验揭示了三个关键观察。首先,CKA和SVCCA在整个训练过程中逐渐降低,表明Transformer层之间的表征专业化增强。其次,内在维度在稳定之前持续增加,表明表征流形逐渐扩展为更大的局部可访问自由度集合。第三,token协方差和耦合分析表明,强大的token交互结构在整个训练过程中持续存在,挑战了表征复杂性增加主要源于token渐进独立的假设。这些发现表明,表征复杂性和层专业化在训练过程中同时出现。流形扩展似乎在没有token解耦的情况下发生。总之,这些观察提出一个新假设,即Vision Transformers在学习过程中通过逐渐更丰富的变换增加表征复杂性,同时保持强大的token交互结构。
英文摘要:
While Vision Transformers have achieved remarkable success across computer vision and language applications, the geometric evolution of their internal representations throughout training remains insufficiently understood. Existing analyses primarily focus on attention mechanisms and downstream performance, leaving the evolution of representation geometry largely unexplored. In this work, we present Transformer Geometry Observatory-II (TGO-II), a representation geometry analysis framework designed to investigate how Transformer representations evolve during supervised training. TGO-II analyzes Vision Transformer (ViT-Small/16) representations using Centered Kernel Alignment (CKA), Singular Vector Canonical Correlation Analysis (SVCCA), Two-Nearest Neighbor Intrinsic Dimensionality (TwoNN-ID), and token covariance analysis. Our experiments reveal three key observations. First, both CKA and SVCCA progressively decrease throughout training, indicating increasing representational specialization across Transformer layers. Second, intrinsic dimensionality consistently increases before stabilizing, suggesting progressive expansion of the representation manifold into a larger set of locally accessible degrees of freedom. Third, token covariance and coupling analyses demonstrate that strong token interaction structure persists throughout training, challenging the hypothesis that increasing representational complexity arises primarily from progressive token independence. These findings suggest that representation complexity and layer specialization emerge simultaneously during training. Manifold expansion appears to occur without token decoupling. Together, these observations motivate a new hypothesis in which Vision Transformers increase representational complexity through progressively richer transformations while preserving strong token interaction structure during learning.