arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越线性概念:发现并对齐大型语言模型中的非线性概念流形

Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff

arXiv 2610.01821首次发表:更新:

发表机构

University of Groningen; Carl von Ossietzky University of Oldenburg(格罗宁根大学; 奥尔登堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将非线性多维概念发现适配至LLM激活,提出概念对齐分数,发现跨层、跨模型及训练阶段的概念流形结构,揭示句法到语义的转变及训练依赖的多语言共享。

AI 中文摘要

理解大型语言模型(LLM)中的信息处理过程,需要剖析其内部令牌表示的几何组织。现有的机械可解释性(MI)方法试图提取概念,但受限于强线性假设,而这一假设正受到非线性特征流形证据的挑战。我们通过将非线性多维概念发现(NLMCD)从计算机视觉领域适配到令牌级LLM激活,从而超越线性概念,将概念建模为低维流形。为了跨层和跨模型比较概念流形,我们引入了一种基于概念的对齐(CBA)分数,这是一种广义Rand指数,无需显式特征匹配即可衡量几何邻近性。我们的分析得出六个关键发现:(i)相邻层合理性检查显示,CBA比基于PCA或CKA的线性基线更敏感;(ii)逐层对齐矩阵揭示了中间层和后期层中的两个块结构,这些结构在不同模型中保持一致,但被线性指标所掩盖;(iii)概念组合在网络的绝大部分阶段仍以句法为主导,之后在后期层逐渐转向混合的句法-语义概念,并向最终层增加输出导向性;(iv)英语和普通话之间的多语言概念共享是训练依赖性的而非普遍性的,在Qwen中最强,在Llama中较弱,在GPT-2中缺失;(v)模型间对齐反映了这一结构,同族不同规模的Qwen模型之间对应性强,而跨模型族的对齐较弱;(vi)在Tulu-3训练阶段中,相邻阶段之间的对齐最高,基础模型与SFT之间的变化最大,而后续的偏好对齐阶段(DPO、RLVR)对早期层影响不大,且RLVR在后期层中大多保留了DPO的概念。

英文摘要

Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of non-linear feature manifolds. We move beyond linear concepts by adapting Non-Linear Multi-Dimensional Concept Discovery (NLMCD) from computer vision to token-level LLM activations, modeling concepts as low-dimensional manifolds. To compare concept manifolds across layers and models, we introduce a concept-based alignment (CBA) score, a generalized Rand index that measures geometric proximity without explicit feature matching. Our analysis yields six key findings: (i) a neighboring-layer sanity check shows CBA is more sensitive than PCA- or CKA-based linear baselines; (ii) layer-by-layer alignment matrices reveal two block structures in intermediate and late layers, consistent across models and obscured by linear metrics; (iii) concept composition remains syntax-dominated through most of the network before giving way to increasingly mixed syntactic-semantic concepts in later layers, with increasing output-orientation toward the final layers; (iv) multilingual concept sharing between English and Mandarin is training-dependent rather than universal, strongest in Qwen, weaker in Llama, and absent in GPT-2; (v) inter-model alignment mirrors this structure, with strong correspondence between same-family Qwen models of different scale but weak alignment across model families; and (vi) across Tulu-3 training stages, alignment is highest between adjacent stages, with the largest shift between the base model and SFT, while subsequent preference-alignment stages (DPO, RLVR) leave early layers largely unchanged and RLVR mostly preserves DPO's concepts in late layers.

Comments24 pages, 13 figures. Code: https://anonymous.4open.science/r/NLMCD-NLP-C5E7

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑