arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05575cs.LGcs.CV

胶囊透镜:在模型表示中定位与追踪概念几何结构

Capsule Lens: Locating and Tracking Concept Geometry in Model Representations

Yiming Tang, Harshvardhan Saini, Samyak Jha, Huaming Chen, Xufeng Duan, Dianbo Liu

首次发表
浏览论文内容

中文总结 AI 辅助

胶囊透镜通过将概念区域拟合为可解释的胶囊几何体,在静态和动态表示中定位并追踪概念几何,揭示不同训练设置下的表示漂移模式。

中文摘要 AI 辅助

理解概念如何编码在机器学习模型的内部表示中,是机制可解释性的核心问题,对于深度学习科学以及日益强大模型的可靠部署都至关重要。现有的解释模型表示的方法主要将表示映射到更可解释的空间,并不直接刻画概念如何占据表示空间;已有多种假设被提出,但往往缺乏严格验证,且大多聚焦于静态表示。在本工作中,我们提出了胶囊透镜(Capsule Lens),一个将概念所占据的区域与一种简单、可追踪的几何形式——胶囊(capsule)相匹配的框架,该胶囊由若干可解释参数定义,以闭式形式拟合每个概念的几何结构,并在留出样本上进行验证。我们在两种主要场景中应用胶囊透镜:静态表示和动态表示。在静态表示上,我们展示了如何在不同模型中定位概念几何结构,以及跨度曲线和范数曲线如何揭示重要的几何特征。在动态表示上,我们呈现了三个案例研究,追踪由不同训练设置引起的表示漂移:CLIP预训练、视觉问答上的强化学习后训练,以及数学推理上的强化学习后训练。这些分析揭示了性质不同的几何动态,范围从CLIP预训练中的广泛网络级重构到强化学习后训练中的局部且概念特定的变化。我们的结果包括与现有文献一致的发现以及新颖的观察。我们相信胶囊透镜是定位、分析和追踪静态与动态表示中概念几何结构的一个有前景的工具。

英文摘要

Understanding how concepts are encoded in the internal representations of machine learning models is a central problem in mechanistic interpretability, essential both for the science of deep learning and for the trustworthy deployment of increasingly capable models. Existing approaches to interpret model representations mainly map representations onto more interpretable spaces and do not directly characterize how concepts occupy representation space; various hypotheses have been proposed, but often lack of rigorous validation and largely focus on static representations. In this work, we introduce Capsule Lens, a framework that matches the region a concept occupies with a simple, trackable geometric form, a capsule, defined by several interpretable parameters, fitted in closed form to each concept's geometry and validated on held-out samples. We apply Capsule Lens in two major settings: static and dynamic representations. On static representations, we demonstrate how to locate concept geometry across various models, and how the span and norm curves uncover important geometric characteristics. On dynamic representations, we present three case studies tracking representation drifts induced by distinct training settings, CLIP pretraining, RL post-training on visual question answering, and RL post-training on mathematical reasoning. These analyses reveal qualitatively different geometric dynamics, ranging from broad network-wide restructuring in CLIP pretraining to localized and concept-specific changes in RL post-training. Our results include findings aligned with existing literature as well as novel observations. We believe Capsule Lens stands as a promising tool for locating, analyzing, and tracking concept geometry in both static and dynamic representations.

发表机构

  • National University of Singapore(新加坡国立大学)
  • Indian Institute of Technology (ISM) Dhanbad(印度理工学院(ISM)丹巴德分校)
  • The University of Sydney(悉尼大学)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

↑