语言模型表示中几何复杂度的局部与全局 regime
Local and Global Regimes of Geometric Complexity in Language Model Representations
- Universitat Pompeu Fabra (UPF)(庞培法布拉大学(UPF))
- ICREA(加泰罗尼亚研究与高级研究所(ICREA))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究探讨词汇多样性对语言模型本征维度(ID)估计的影响,发现其存在尺度依赖的局部与全局 regime 转变,推导了反转点公式,为LLMs内部流形结构研究提供了新视角。
AI中文摘要:
本征维度(ID)被广泛用于探测语言模型的表示复杂度,但目前仍不清楚ID的差异是反映语言本身的属性还是底层数据集构建方式的人工产物。本文专门关注数据集中的词汇多样性(即唯一末-token 项的数量)如何影响该数据集的ID估计。我们发现两种 regime 之间存在尺度依赖的转变:在低词汇多样性下,唯一末词较少的条件会产生更高的ID;而在高词汇多样性下,该顺序反转,唯一词更多的条件会产生更高的ID。我们推导了一个精确的、无参数的公式来计算这种反转发生的点,该公式与所有测试尺度下观察到的转变点都匹配。一方面,我们的结果强调,在将一组表示的本征维度解释为其复杂度的直接线索时必须谨慎;另一方面,我们对两种ID regime 的发现揭示了大型语言模型(LLMs)中语言数据组织的一般原则,为其内部流形结构提供了新的见解。
英文摘要:
Intrinsic dimensionality (ID) is widely used to probe the representational complexity of language models, but it remains unclear whether ID differences reflect properties of language itself or artefacts of how the underlying dataset was constructed. In this paper, we focus specifically on how lexical diversity, the number of unique last-token items present in a dataset, affects ID estimates of that dataset. We find a scale-dependent transition between two regimes: at low lexical diversity, conditions with fewer unique final words produce higher ID, while at high lexical diversity, this ordering reverses, and conditions with more unique words produce higher ID. We derive an exact, parameter-free formula for the point at which this reversal occurs, which matches the observed transition point at every scale tested. On the one hand, our results highlight how care must be taken when interpreting the intrinsic dimensionality of a set of representations as a straightforward cue of their complexity. On the other hand, our discovery of the two ID regimes reveals a general principle of organisation of linguistic data in LLMs that sheds new light on their inner manifold structures.