AI 中文总结
本研究探究Transformer各层语法角色对表征流形几何的影响,分析ID演化、邻域结构变化,对比编码器与解码器的差异,发现几何特征可恢复token语法角色并用于语义演化解释。
AI 中文摘要
Transformer表征通过高维向量空间中的轨迹来描述,这些轨迹会随 token 在各层整合关系上下文而动态变化。这类数据往往集中在低维子流形上,这种压缩形式由固有维度(Intrinsic Dimensionality, ID)量化,即无需显著信息损失即可表征数据所需的最小独立变量数。本研究探究 token 的语法角色(由其词性(Part-of-Speech, PoS)标记标注)是否会影响该流形的局部几何形态,为此开展四项工作:(1)探究 ID 的逐层演化,发现封闭类词项比开放类词项更早扩张且更早收缩;(2)表明 ID 的扩张与收缩可由邻域结构的变化解释,进而可由句子内单词间关系的变化解释;(3)对比编码器(ModernBERT、bigbird-roberta-large)与解码器(gemma-2-2B、Llama-3.2-3B),发现两类模型在各层的演化方式不同,与各自整合上下文的方式一致;(4)表明仅几何特征即可恢复 token 的语法角色,并将其用于解释下游分类任务中各词性的语义内容如何随各层演化。
英文摘要
Transformer representations describe trajectories through high-dimensional vector spaces, which are shaped dynamically as tokens incorporate relational context across layers. Such data tend to concentrate on lower-dimensional sub-manifolds, a form of compression quantified by the Intrinsic Dimensionality (ID), the minimum number of independent variables needed to represent them without significant information loss. In this work, we ask whether the grammatical role of tokens, as marked by their part-of-speech (PoS) tag, shapes the local geometry of this manifold. To this end: (1) We investigate the layer-wise evolution of ID, finding that closed-class items expand earlier and collapse sooner than open-class ones; (2) We show its expansion and contraction to be explained by changes in the neighborhood structure, and hence in the relations between words within a sentence; (3) We compare encoders (ModernBERT, bigbird-roberta-large) and decoders (gemma-2-2B, Llama-3.2-3B), finding that the two families evolve differently across layers, consistently with how each integrates context;(4) We show that geometric features alone recover a token's grammatical role, and use them to interpret how the semantic content of each PoS evolves across layers in a downstream classification task.