IRIS:一种受视觉皮层启发的用于分析视觉Transformer中方向选择性的框架
IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers
浏览论文内容
中文总结 AI 辅助
该研究提出受视觉皮层启发的IRIS框架,通过RSS等神经科学指标分析ViT中方向选择性的形成,发现训练范式决定方向选择性、层的方向选择性变化规律,且指标可指导解冻层数以优化下游泛化。
中文摘要 AI 辅助
视觉Transformer(ViTs)已成为许多感知任务中图像编码的事实上的标准。尽管它们在经验上取得了成功,但由于缺乏归纳偏置(ViTs全局处理信息而非依赖局部结构),其如何编码低级特征的机制仍不清楚。相比之下,生物视觉系统通过结合视野中小的局部区域的信息来构建低级特征,例如初级视觉皮层中的方向选择性,这些特征是通用表示,在多个专门的神经通路中共享且被需要,不同于更高层次的、特定于任务的语义特征。这就提出了一个问题:这种基于生物学的特征是否会在ViTs中出现。在这项工作中,我们通过引入一套受神经科学启发的指标:表示相似性得分(RSS)、方向招募得分(ORS)和方向调谐带宽,来系统研究ViTs中方向选择性如何出现,以量化方向如何在表示几何中以及作为模型深度的函数进行编码。通过广泛分析,我们发现:(1)训练范式是方向选择性的最强决定因素,具有相同目标的模型无论规模如何,都会在相当的相对深度达到峰值;(2)许多单元在训练早期就具有方向选择性,中低层会随时间招募更多此类单元,而更深层会失去选择性并拓宽调谐以转向语义编码;(3)我们的指标为应解冻多少层以实现最佳下游泛化提供了一种机制启发式方法。我们的框架提供了一种在ViT训练过程中跟踪基于生物学的特征的方法,探究了所需属性如何在Transformer表示中编码,并建立了对ViTs如何跨任务泛化的系统理解。
英文摘要
Vision transformers (ViTs) have become the de facto standard for image encoding across many perception tasks. Despite their empirical success, it remains mechanistically unclear how they encode low-level features, given their lack of inductive biases: ViTs process information globally rather than relying on local structure. Biological visual systems, in contrast, build low-level features, such as orientation selectivity in the primary visual cortex, by combining information from small, localized regions of the visual field. These features are general-purpose representations, shared and required across multiple specialized neural pathways, unlike higher-level, task-specific semantic features. This raises the question if such biologically-grounded features arise in ViTs. In this work, we systematically study how orientation selectivity emerges in ViTs by introducing a suite of neuroscience-inspired metrics: representational similarity score (RSS), orientation recruitment score (ORS), and orientation tuning bandwidth to quantify how orientation is encoded in representational geometry and as a function of model depth. Through extensive analysis, we find that: (1) the training paradigm is the strongest determinant of orientation selectivity, with models sharing an objective, peaking at comparable relative depths regardless of scale (2) many units are orientation-selective early in training, with early-to-middle layers recruiting more such units over time, while deeper layers lose selectivity and broaden their tuning toward semantic encoding and (3) our metrics offer a mechanistic heuristic for how many layers to unfreeze for best downstream generalization. Our framework presents a way to track biologically-grounded features during ViT training, probes how desired properties are encoded in transformer representations, and builds a systematic understanding of how ViTs generalize across tasks.
发表机构
- University of Washington(华盛顿大学)
- Microsoft(微软公司)
- TU Munich(慕尼黑工业大学)
- Gladstone Institute(格拉德斯通研究所)
- Nvidia(英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。