arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你是否和我想的一样?:探究神经架构中的概念分离

Are You Thinking What I am Thinking? : Examining Conceptual Separation in Neural Architectures

Jaee Ponde, Roshni Agarwal, Subhashis Banerjee

arXiv 2609.00764首次发表:更新:

发表机构

Ashoka University; Truth Audit Labs; Karya AI(阿肖克大学; 真相审计实验室; 卡尔亚人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过分析CNN与LLM的内部激活,探究概念分离现象,发现其能揭示输出准确率无法体现的模型表征结构,可用于诊断模型的概念表征鲁棒性。

AI 中文摘要

神经网络越来越多地被用于识别定义明确及模糊的概念,但输出层面的指标几乎无法揭示这些概念是如何被内部表征的。本研究探究这些网络是否表现出「概念分离」:即同一概念的样本是否形成连贯的表征,以及相关概念在表征空间中是否靠得更近。我们通过对内部激活值的几何与分布分析,探究卷积神经网络(CNN)与大型语言模型(LLM)中的这种概念组织。在CNN中,熟悉的ImageNet概念形成连贯且语义有序的表征,而对于未见过的概念,这种连贯性会减弱,且在类内域偏移中受损;在LLM中,明显不同的领域仍保持良好分离,相关子领域靠得更近,模糊主题的区分在均值和协方差层面均崩溃。这些结果表明,概念分离可揭示仅靠输出准确率无法发现的结构,且可作为诊断模型对所识别概念的表征鲁棒性的有用工具。代码和数据可在GitHub获取。

英文摘要

Neural networks are increasingly employed to identify both well-defined and ambiguous concepts, yet output-level metrics reveal little about how those concepts are represented internally. Our study asks if these networks exhibit \textit{conceptual separation}: if examples of the same concept form coherent representations, and whether related concepts lie closer together in the representation space. We examine this conceptual organisation in Convolutional Neural Networks (CNNs) and Large Language Models (LLMs) through geometric and distributional analysis of their internal activations. In CNNs, familiar ImageNet concepts form coherent and semantically ordered representations, while this coherence weakens for unseen concepts and suffers within-class domain shift. In LLMs, clearly distinct domains remain well separated, related subdomains move closer together, and the distinction between ambiguous topics collapses at both the mean and covariance level. These results suggest that conceptual separation can reveal structure that output accuracy alone cannot, and may serve as a useful diagnostic of how robustly a model represents the concepts it is asked to identify. Code and data available on \href{https://github.com/JaeeRoshniCapstoneProject/Are-You-Thinking-What-I-m-Thinking-Examining-Conceptual-Separation-in-Neural-Architectures}{GitHub}.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑