发表机构
University of Bologna; ARCES; DEI(博洛尼亚大学; 先进计算机科学与工程研究中心; 电子、信息与生物工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出基于SVD的典型性图方法,推导PAS与MLSV分数,在未重新训练或接触OOD数据的情况下,于ViT-B/16微调CIFAR-100任务上实现了有竞争力的分布外检测性能。
AI 中文摘要
我们提出一种利用视觉Transformer(ViTs)学习参数几何结构分析其内部表征的方法。对每个仿射层的权重矩阵进行奇异值分解(SVD),并将激活值投影到前导右奇异向量以获得紧凑的层内表征。随后在每一层拟合类条件密度模型,生成逐类典型性分数,将其沿深度堆叠得到典型性图,即类特定证据在网络中演化的二维摘要。从这些图中,我们推导了两个用于分布外(OOD)检测的事后分数:原型对齐分数(PAS),用于衡量与类参考原型模式的一致性;以及多层软投票(MLSV)分数,无需存储原型即可捕捉跨层共识。在针对CIFAR-100微调的ViT-B/16上,所提出的分数在无需重新训练或接触OOD数据的情况下,实现了具有竞争力的检测性能。
英文摘要
We present a method for analyzing the internal representations of Vision Transformers (ViTs) exploiting the geometry of their learned parameters. Each affine layer's weight matrix is factored via Singular Value Decomposition (SVD), and activations are projected onto the leading right singular vectors to obtain compact, layer-intrinsic representations. A class-conditional density model is then fitted at each layer, producing per-class \emph{typicality scores} that are stacked across depth into \emph{typicality maps}: two-dimensional summaries of how class-specific evidence evolves through the network. From these maps, we derive two post-hoc scores for Out-Of-Distribution (OOD) detection: a \emph{Prototype Alignment Score} (PAS), measuring agreement with class reference prototype patterns, and a \emph{Multi-Layer Soft Voting} (MLSV) score, capturing cross-layer consensus without stored prototypes. On ViT-B/16 fine-tuned on CIFAR-100, the proposed scores achieve competitive detection performance without retraining or OOD exposure.