垂直融合:压缩内部表示以实现稳健的视觉Transformer分类
Vertical Fusion: Condensing Internal Representations for Robust ViT Classification
浏览论文内容
中文总结 AI 辅助
研究挑战视觉Transformer仅用最后一层做下游任务的传统,提出可恢复性概念。通过实验发现中间层能纠正部分最后层误分类样本,增益源于冗余 - 正确性对应。提出VFusion垂直聚合策略,该策略在多方面优于基线,为水平集成方法提供稳健高效替代。
中文摘要 AI 辅助
尽管视觉Transformer(ViT)展现出丰富的中间表示,但几乎仅被用作黑盒特征提取器,仅考虑最后一层用于下游任务。我们引入可恢复性概念对此传统做法提出挑战,即中间表示纠正最后一层故障的能力。通过在16个数据集的每个模型深度评估独立分类探测器,发现中间探测器能正确分类最后一层探测器误分类的18%至76%的样本。这些增益并非主要由预测多样性驱动,而是由冗余 - 正确性对应关系导致。虽然现有水平集成策略可提高性能,但计算成本高且忽略单个模型内的垂直信号。为弥合这一差距,我们提出VFusion,一种有原则的垂直聚合策略,通过可学习映射到低维潜在空间来合成跨ViT内部层次结构的特征。VFusion在分布内和分布外设置中均显著优于现有聚合基线,缩小了最佳单个层与理论最优性能之间45%的准确率差距,且在不同模型大小和预训练模式下均有增益,证明其为水平集成方法提供了稳健且高效的替代方案。代码可通过此https链接获取。
英文摘要
Despite exposing rich intermediate representations, Vision Transformers (ViTs) are almost exclusively utilized as black-box feature extractors, where only the last layer is considered for downstream tasks. We challenge this convention by introducing the notion of recoverability: the capacity of intermediate representations to correct last-layer failures. By evaluating independent classification probes at every model depth across 16 datasets, we observe that intermediate probes correctly classify 18% to 76% of samples that the last-layer probe misclassifies. We show that these gains are not primarily driven by predictive diversity, but by a redundancy-correctness correspondence, where the internal hierarchy acts as a series of stable, redundant probes of a shared discriminative signal. While established horizontal ensemble strategies (i.e., across multiple models) can improve performance, they incur high computational cost and ignore this vertical signal within a single model. To bridge this gap, we propose VFusion, a principled vertical aggregation strategy employing a learnable mapping into a low-dimensional latent space that synthesizes features across the internal ViT hierarchy. VFusion substantially outperforms established aggregation baselines in both in-distribution and out-of-distribution settings, notably closing 45% of the accuracy gap between the best individual layer and a theoretical oracle performance. Our gains consistently generalize across model sizes and pre-training regimes, confirming that VFusion offers a robust and efficient alternative to horizontal ensemble methods. The code is available at https://github.com/francescodisalvo05/vit-vertical-fusion.
发表机构
- University of Bamberg(班贝格大学)
- AIML, The University of Adelaide(阿德莱德大学AIML)
机构由 AI 辅助整理,请以论文原文为准。