发表机构
Hebrew University; The Racah Institute of Physics; Edmond and Lily Safra Center for Brain Sciences; Cornell Tech, Cornell University; Center for Brain Science, Harvard University(希伯来大学; 拉卡物理研究所; 埃德蒙和莉莉·萨夫拉脑科学中心; 康奈尔大学科技校区; 哈佛大学脑科学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究揭示跨模态AI模型分类任务存在通用表征几何,推导平均场理论准确预测分类准确率,发现分类由稀疏质心对齐结构支配,关联稀疏特征提取方法。
AI 中文摘要
神经表征已成为研究现代AI模型内部机制的核心工具,但其复杂的高维结构使其难以解释。我们表明,分类任务会产生一种通用的表征几何,该几何在视觉、音频和语言处理领域的最先进模型中是共享的。关键结构在于,类内变异性在表征空间中并非随机,相反,其与分类器相关的成分与该类自身的质心及其竞争类的质心具有强且结构化的相关性。基于这一观察,我们推导了一种解析平均场理论,该理论主要由真实类和竞争类质心坐标的变异性,以及补偿真实表征非高斯统计的类半径全局归一化所支配。该理论能准确预测不同架构和模态下的分类准确率,相关几何量随模型规模系统提升,与观察到的准确率提升相呼应。该理论的一个显著特征是其稀疏性:准确预测仅需少量与真实类及其最强竞争对手相关的质心坐标,这将我们的框架与稀疏特征提取方法(如稀疏自编码器)联系起来。这些结果共同为神经表征提供了一种简约的预测理论,并表明深度网络中的分类由嵌入在完整高维表征空间内的稀疏、与质心对齐的结构所支配。
英文摘要
Neural representations have become a central tool for studying the internal mechanisms of modern AI models, yet their complex high-dimensional structure makes them difficult to interpret. We show that classification tasks give rise to a universal representational geometry, shared across state-of-the-art models in vision, audio, and language processing. The key structure is that within-class variability is not random in representation space. Instead, its classifier-relevant component has strong and structured correlations with the class's own centroid and with the centroids of its competing classes. Building on this observation, we derive an analytical mean-field theory governed mainly by the variability along true-class and rival-class centroid coordinates, together with a global renormalization of the class radius that compensates for the non-Gaussian statistics of real representations. The theory accurately predicts classification accuracy across architectures and modalities. The relevant geometric quantities improve systematically with model scale, mirroring the observed gains in accuracy. A striking feature of the theory is its sparsity: accurate prediction requires only a small set of centroid coordinates associated with the true class and its strongest rivals - connecting our framework to sparse-feature extraction approaches such as sparse autoencoders. Together, these results provide a parsimonious predictive theory of neural representations and suggest that classification in deep networks is governed by a sparse, centroid-aligned structure embedded within the full high-dimensional representation space.
Comments33 pages, 13 figures, 14 tables