arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

短行中心簇会放大词元嵌入的内在维度估计值

A Hub of Short Rows Inflates Intrinsic Dimension Estimation of Token Embeddings

Alexandre Quemy

arXiv 2608.29702首次发表:更新:

发表机构

Hother Labs(霍瑟实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现词元嵌入表的原点附近短行中心簇会使TwoNN等内在维度估计值偏高,移除该簇后11个模型的维度会收缩,归一化行也可得到较低读数,复现验证了Pythia的相关实验结果。

AI 中文摘要

词元嵌入表在原点附近存在一个短行中心簇,我们证明该簇会对近邻内在维度(ID)估计器的输出产生偏差。由于测度集中性,一个词元到中心簇的距离比到其他任何词元都近,因此其前两个近邻均为中心簇行,且距离几乎相等。结果,TwoNN等ID估计器会返回远高于真实ID的维度值。按单个词元测量时,维度呈重尾分布;按整个词汇表测量时,维度随模型参数数量增长。但当移除该中心簇后,重尾消失,11个模型(从GPT-2到K3、GLM-4.7等模型)的测量维度均收缩至狭窄范围。该中心簇如同一个开关:仅需数百行即可完全放大估计值。我们复现了一项实验,该实验表明Pythia的词元嵌入表的内在维度(ID)随参数数量增长,在160M到12B参数间从27升至122。我们证明,移除中心簇后该结果消失:各参数规模下嵌入表的ID均为10至17。该中心簇包含欠训练词元检测器标记的子集,但在Pythia上,我们检测并整体移除的中心簇在训练过程中被更新:这些行的特征仅为长度,而非未被更新。最后,我们证明,对行进行归一化而非移除,会得到相同的较低维度读数。

英文摘要

A token-embedding table holds a hub of short rows near its origin, and we show that this cluster biases what nearest-neighbor intrinsic-dimension (ID) estimators report. Because of the concentration of measure, a token is closer to the central cluster than to any other token, so its first two neighbors are both hub rows at nearly the same distance. As a result, the ID estimators such as TwoNN return a dimension far above the real ID. Measured one token at a time, dimension is a heavy-tailed distribution. Measured on the full vocabulary, it grows with the model's parameter count. However, when we remove the hub, the heavy tail disappears and the measured dimension collapses to a narrow range for eleven models, from GPT-2 to models such as K3 and GLM-4.7. The hub acts as a switch: a few hundred rows are enough to fully inflate the estimate. We reproduced an experiment stating that the intrinsic dimension (ID) of Pythia's token-embedding table grows with the parameter count, from $27$ to $122$ between 160M and 12B parameters. We show that this result disappears when the hub is removed: the table then reads $10$ to $17$ at every size. The hub contains a subset of the population that under-trained-token detectors flag, but on Pythia the hub that we detected and removed as a whole was updated during training: what seem to characterize these rows is simply their length, not an absence of updates. Finally, we show that normalizing the rows instead of removing them gives the same lower reading.

CommentsSubmitted to NeurReps Workshop @ NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑