发表机构
University of Delaware; University of Virginia(特拉华大学; 弗吉尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出用跨层转码器从视觉Transformer中读取概念电路,以揭示其内部世界知识,并通过全局与实例电路实现虚假相关性发现、移除及模型比较,在Waterbird上提升11.0%。
AI 中文摘要
视觉Transformer(ViTs)在跨视觉领域取得了显著的泛化能力,然而关于它们如何在内部表示世界结构,我们知之甚少。为弥补这一空白,我们使用跨层转码器(CLTs)从ViTs中读取概念电路:这是一种有向图,其节点对应稀疏、可解释的概念,边则捕捉跨层的概念交互。我们的方法提供了模型行为的两种互补视角。全局概念电路与输入无关,可直接从学习到的跨层权重中恢复,揭示了模型中编码的可复用“世界知识”。实例概念电路则依赖于输入,识别出特定预测实际使用的概念和路径,从而支持忠实的示例级解释。我们通过三种方式展示了概念电路的有效性:(1)自动虚假相关性发现:利用全局概念电路的统计信息来识别模型内部的捷径依赖。(2)虚假相关性移除:对实例概念电路进行干预,引导模型做出正确预测。实验结果表明,在Waterbird数据集上,我们的方法比现有方法高出11.0%。(3)模型比较:对比不同基础模型(如CLIP与DINO)的全局概念电路,揭示监督范式如何塑造表征结构。我们的代码可在以下网址获取:此https URL。
英文摘要
Vision transformers (ViTs) have achieved remarkable generalization across visual domains, yet little is known about how they internally represent the structure of the world. To address this gap, we use Cross-Layer Transcoders (CLTs) to read concept circuits from ViTs: directed graphs whose nodes correspond to sparse, interpretable concepts and edges capture concept interactions across layers. Our method yields two complementary views of model behavior. The global concept circuit is input-invariant and can be recovered directly from learned cross-layer weights, exposing the reusable "world knowledge" encoded in the model. The instance concept circuit is input-dependent and identifies the concepts and pathways actually used for a specific prediction, enabling faithful example-level explanations. We demonstrate the utility of concept circuits in three ways: (1) Automatic spurious correlation discovery: leveraging the statistics of our global concept circuits to identify shortcut dependencies within the model. (2) Spurious correlation removal: intervening on the instance concept circuit to steer the model towards correct predictions. Empirical results show that our method outperforms existing counterparts by 11.0% on the Waterbird dataset. (3) Model comparison: contrasting the global concept circuits of different foundation models (e.g., CLIP vs. DINO) to reveal how supervision paradigms shape representational structure. Our code is available at https://github.com/deep-real/VisionCLT
CommentsECCV 2026