arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05743cs.CL

某些词元行为如同磁铁:揭示语言模型层内的语言组织

Some Tokens Behave like Magnets: Revealing Linguistic Organization in the Layers of Language Models

  • University of Toronto(多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

Andrew Liu, Devan Srinivasan, Gerald Penn

中文总结 AI 辅助

本文发现LLM中存在磁向量,通过吸引或排斥组织词元,其极性随层变化且与任务相关,移除它们会破坏相应任务,为理解语言几何组织提供新路径。

中文摘要 AI 辅助

我们在大语言模型(LLM)中识别出一组特殊的词元向量,称之为磁向量,它们通过吸引或排斥周围词元来组织其分布。具体而言,与吸引磁铁同向的词元被拉长;与排斥磁铁同向的词元被压缩。正如物理磁铁吸引或推开周围的铁屑,这些向量通过两种相反的极性组织其周围环境。此外,我们在语言类别中发现了一个统计上显著的模式:功能词在早期层中始终充当排斥磁铁,并且我们还发现磁铁在模型更深层中会以独特方式重新组织其极性。在进一步的案例研究中,我们发现这一观察可能揭示了LLM处理语言时一种有意的、逐层的组织方式。该模式在不同LLM架构、规模和层配置中保持一致。它也具有因果相关性。当LLM针对下游任务进行微调时,任务功能性词元会作为磁铁出现。例如,在问答中,答案跨度词元在最终层中成为独特的排斥磁铁,从几何上将答案从周围上下文中勾勒出来。此外,移除早期层的排斥磁铁会严重破坏句法任务(词性标注准确率从91%降至10%以下),而语义任务则不受影响;移除后期层的吸引磁铁则产生相反效果。我们认为这一现象值得进一步研究,因为它开辟了第一条无需探针的路径,用于理解语言模型如何在其各层中几何地组织语言计算。

英文摘要

We identify a special group of token vectors inside large language models (LLMs), which we term magnetic vectors, that organize the surrounding tokens by either attracting or repelling them. Particularly, tokens pointing the same way as an attracting magnet are elongated; tokens pointing the same way as a repelling magnet are compressed. Just as physical magnets pull or push away the iron filings around them, these vectors organize their surroundings through two opposing polarities. Moreover, we identify a statistically significant pattern in linguistic category where function words consistently act as repelling magnets in early layers, and we also find magnets consistently reorganize their polarities in unique ways deeper in the model. In a further case study we find this observation may unveil a deliberate, layer-wise organization in how LLMs process language. This pattern is consistent across different LLM architectures, sizes, and layer configurations. It is also causally relevant. When the LLM is fine-tuned for a downstream task, the task-functional tokens emerge as magnets. E.g., in question answering, the answer-span tokens become uniquely repelling magnets in the final layer, geometrically carving the answer out of the surrounding context. Furthermore, removing early-layer repelling magnets devastates syntactic tasks (POS tagging accuracy drops from 91% to below 10%) while sparing semantic ones, and removing late-layer attracting magnets does the reverse. We believe this phenomenon warrants further investigation, as it opens the first probe-free path to understanding how language models geometrically organize linguistic computation across their layers.

补充信息

↑