arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19887cs.CL

Transformer模型中的词汇抽象泛化:功能词案例研究

Generalization through Lexical Abstraction in Transformer Models: The Case of Functional Words

发表机构Idiap 研究所 · 日内瓦大学
查看机构详情
  • Idiap Research Institute(Idiap 研究所)
  • University of Geneva(日内瓦大学)

机构由 AI 辅助整理,请以论文原文为准。

Giuseppe Samo, Vivi Nastase, Paola Merlo

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探究预训练Transformer模型如何编码功能词及其在平行句中的泛化能力,发现混合训练可揭示共享句法语义结构。

中文摘要 AI 辅助

代词、副词及其他功能词(如they、her、somewhere、there)在语言中常被用来替代具体的名词或短语,当它们的属性(如性别、语法数)为给定语境提供了足够信息时。预训练Transformer模型是否以允许它们像人类一样使用的方式来编码这些功能词?语言模型能否识别诸如“The researchers wrote the paper”和“They wrote it”这类依赖词汇抽象的句子的句法和语义平行性?我们将这些语言学问题映射到预训练Transformer模型的嵌入空间,并比较名词的表征与可替代这些名词的代词和副词的表征,分别在孤立语境和平行词汇化与功能化句子中。随后,我们探测平行词汇化和功能化句子嵌入中共享的句法和语义结构。我们发现,功能词在嵌入空间中位于名词的中心区域,但又有别于名词,这与它们作为各种语境中占位符的行为一致。对平行(词汇化和功能化)句子嵌入的分析显示,它们占据嵌入空间的不同子空间。蒸馏句子结构信息的实验表明,仅用任一类型数据训练都无法揭示共享结构——原因在于功能数据的词汇过度一致性,以及词汇化版本的数据多样性过大。然而,使用功能化与词汇化句子的混合数据进行训练时,共享结构得以浮现。

英文摘要

Pronouns, adverbs and other functional words (such as they, her, somewhere, there) are often used in language to replace concrete nouns or phrases, when their properties - such as gender, grammatical number - provide sufficient information for the given context. Do pretrained transformer models encode such functional words in a manner that allows them to be used like humans do? Can language models recognize the syntactic and semantic parallelism of sentences such as "The researchers wrote the paper" and "They wrote it", which relies on such lexical abstraction? We map these linguistic questions into the embedding space of a pretrained transformer model, and compare representations of nouns, with the representations of the pronouns and adverbs that can replace these nouns, in isolation and in parallel lexicalized and functional sentences. We then probe for shared syntactic and semantic structure in the embeddings of parallel lexicalized and functional sentences. We find that functional words are located centrally compared to nouns, but are also distinct, which is congruent with their behaviour as place-holders in a wide variety of contexts. The analysis of the embeddings of parallel (lexicalized and functional) sentences show them inhabiting different subspaces of the embedding space. Experiments that distil the structural information of the sentence show that training on either type of data does not reveal the shared structure - because of the over-consistency of the vocabulary (in case of the functional data), and the too much variety (in case of the lexicalized versions). However, training with a mix of functional and lexicalized sentences, the shared structure emerges.

补充信息

↑