歧义词对中收敛表示引发的大语言模型发散预测
Divergent large language model predictions from convergent representations in ambiguous word pairs
查看机构详情
- University of Otago(奥塔哥大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究通过分析三种规模的仅解码器Transformer,揭示其解决词汇歧义的机制:中间层表示区分度最高、顶层预测KL散度最大,且顶层表示差异直接决定输出,为嵌入方法的应用提供了启示。
中文摘要 AI 辅助
在本研究中,我们通过逐层分析三种不同参数规模的模型(GPT-2-Small-117M、Llama-3.2-3B、Qwen2.5-32B),探究仅解码器的Transformer如何解决词汇歧义。对于同形异义词和多义词,我们发现表示在中间层达到最大区分度,随后在顶层部分重新收敛,而它们的下一个词预测之间的KL散度在最终层达到最大值。激活修补实验提供了因果证据:尽管在嵌入空间中相似性明显增加,但顶层的表示差异直接决定输出。我们的单层消融实验表明,尽管模型在逐层脆弱性上存在质的差异,仍能实现同等的歧义消解。这些发现为近期观察到的现象提供了机制:尽管模型表现强劲,其内部嵌入相似性与行为输出的相关性却很低。因此,语义区分仍然存在,但对于依赖顶层余弦相似性的嵌入方法(如语义搜索、检索和聚类)而言,这些区分变得越来越不可见。
英文摘要
In this work we investigate how decoder-only transformers resolve lexical ambiguity through layer-by-layer analysis of three models spanning three parameter sizes (GPT-2-Small-117M, Llama-3.2-3B, Qwen2.5-32B). For both homonyms and polysemes, we find that representations become maximally distinct in middle layers, then partially reconverge in late layers, while the KL divergence between their next-token predictions reaches its maximum in the final layers. The activation patching experiment provides causal evidence that late-layer representational differences directly determine outputs despite apparent increased similarity in embedding space. Our single-layer ablation experiment indicates that models achieve equivalent disambiguation despite qualitatively different layer-wise vulnerabilities. These findings offer a mechanism for recent observations where models' internal embedding similarities show low correlation with their behavioural outputs despite strong performance. The semantic distinctions therefore remain present but become increasingly invisible to similarity measures over the embeddings, with implications for embedding-based methods such as semantic search, retrieval, and clustering that rely on late-layer cosine similarity.