AI 中文总结
本文通过稀疏自编码器分析DINOv2中寄存器与高范数离群令牌的功能,发现前者编码高层语义,后者编码低层纹理,且因果消融显示寄存器特征对表示影响显著更大,证实令牌特化现象。
AI 中文摘要
自监督视觉Transformer(如DINOv2)能够学习丰富的视觉表示,但其内部令牌的功能仍未被充分理解。近期架构引入专用寄存器令牌以减少背景区域中出现的高范数离群补丁令牌,然而这两类令牌的语义与功能角色尚未完全确立。本文通过在DINOv2中训练稀疏自编码器(SAEs)来分析寄存器令牌与离群令牌的激活。利用自动化可解释性流程、UMAP聚类及CLIP空间交叉验证,我们发现寄存器令牌特征与高层语义概念关联更强。相比之下,离群令牌特征更常与低层结构、背景及纹理主导模式相关。因果消融进一步揭示了显著的功能不对称性:破坏激活最强的寄存器衍生特征导致表示余弦相似度下降48.17%,而破坏离群衍生特征仅导致0.31%的下降。综上,我们的结果为自监督ViT中的令牌特化提供了证据。
英文摘要
Self-supervised Vision Transformers (ViTs), such as DINOv2, learn rich visual representations, but the functions of their internal tokens remain poorly understood. Recent architectures introduce dedicated register tokens to reduce high-norm out- lier patch tokens that emerge in background re- gions, yet the semantic and functional roles of both token types have not been fully established. In this paper, we analyze these roles by training sparse autoencoders (SAEs) on register-token and outlier-token activations in DINOv2. Using an automated interpretability pipeline, UMAP clus- tering, and CLIP-space cross-checks, we find that register-token features are more strongly associ- ated with high-level semantic concepts. Outlier- token features, by contrast, are more often associ- ated with lower-level structural, background, and texture-dominant patterns. Causal ablations fur- ther reveal a substantial functional asymmetry: disrupting top-activating register-derived features produces a 48.17% drop in representation cosine similarity, whereas disrupting outlier-derived fea- tures produces only a 0.31% drop. Together, our results provide evidence for token specialization in self-supervised ViTs.