arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10391cs.CVcs.LGeess.IV

垂直融合:压缩内部表示以实现稳健的视觉Transformer分类

Vertical Fusion: Condensing Internal Representations for Robust ViT Classification

Francesco Di Salvo, Shyam Nandan Rai, Hamed Damirchi, Ignacio Meza De la Jara, Sebastian Doerrich, Marco Lents, Christian Ledig

首次发表
浏览论文内容

中文总结 AI 辅助

研究挑战视觉Transformer仅用最后一层做下游任务的传统,提出可恢复性概念。通过实验发现中间层能纠正部分最后层误分类样本,增益源于冗余 - 正确性对应。提出VFusion垂直聚合策略,该策略在多方面优于基线,为水平集成方法提供稳健高效替代。

中文摘要 AI 辅助

尽管视觉Transformer(ViT)展现出丰富的中间表示,但几乎仅被用作黑盒特征提取器,仅考虑最后一层用于下游任务。我们引入可恢复性概念对此传统做法提出挑战,即中间表示纠正最后一层故障的能力。通过在16个数据集的每个模型深度评估独立分类探测器,发现中间探测器能正确分类最后一层探测器误分类的18%至76%的样本。这些增益并非主要由预测多样性驱动,而是由冗余 - 正确性对应关系导致。虽然现有水平集成策略可提高性能,但计算成本高且忽略单个模型内的垂直信号。为弥合这一差距,我们提出VFusion,一种有原则的垂直聚合策略,通过可学习映射到低维潜在空间来合成跨ViT内部层次结构的特征。VFusion在分布内和分布外设置中均显著优于现有聚合基线,缩小了最佳单个层与理论最优性能之间45%的准确率差距,且在不同模型大小和预训练模式下均有增益,证明其为水平集成方法提供了稳健且高效的替代方案。代码可通过此https链接获取。

英文摘要

Despite exposing rich intermediate representations, Vision Transformers (ViTs) are almost exclusively utilized as black-box feature extractors, where only the last layer is considered for downstream tasks. We challenge this convention by introducing the notion of recoverability: the capacity of intermediate representations to correct last-layer failures. By evaluating independent classification probes at every model depth across 16 datasets, we observe that intermediate probes correctly classify 18% to 76% of samples that the last-layer probe misclassifies. We show that these gains are not primarily driven by predictive diversity, but by a redundancy-correctness correspondence, where the internal hierarchy acts as a series of stable, redundant probes of a shared discriminative signal. While established horizontal ensemble strategies (i.e., across multiple models) can improve performance, they incur high computational cost and ignore this vertical signal within a single model. To bridge this gap, we propose VFusion, a principled vertical aggregation strategy employing a learnable mapping into a low-dimensional latent space that synthesizes features across the internal ViT hierarchy. VFusion substantially outperforms established aggregation baselines in both in-distribution and out-of-distribution settings, notably closing 45% of the accuracy gap between the best individual layer and a theoretical oracle performance. Our gains consistently generalize across model sizes and pre-training regimes, confirming that VFusion offers a robust and efficient alternative to horizontal ensemble methods. The code is available at https://github.com/francescodisalvo05/vit-vertical-fusion.

发表机构

  • University of Bamberg(班贝格大学)
  • AIML, The University of Adelaide(阿德莱德大学AIML)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑