arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

曝光非必需:语言模型中学习不同类别的并列结构

Exposure is Optional: Learning Unlike Coordination in Language Models

Jiamu Luo, Shane Steinert-Threlkeld

arXiv 2607.20251首次发表:更新:

发表机构

University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究语言模型中不同类并列结构习得是否需直接曝光,用过滤语料库训练GPT-2模型,发现无需直接曝光,模型能泛化处理,还揭示了模型处理方式及可从同类并列结构学习,助力理解语言模型结构表示。

AI 中文摘要

并列结构作为一种基本的语言结构,仍然是激烈辩论的主题,其确切性质仍然困扰着理论语言学。一种常见观点认为只有同类成分才能并列,这一观点受到自然语言中许多不同类并列结构的挑战。我们将语言模型视为计算测试平台,研究不同类并列结构的习得是否需要训练数据中的直接曝光,或者它是否可以从一般的组合能力中有机出现。我们使用过滤语料库训练(FiCT),在去除所有不同类并列结构实例的语料库上训练GPT - 2模型。我们发现直接曝光并非必要:在过滤数据上训练的模型成功地将不同类并列结构进行了泛化,在困惑度和语法判断上与在未过滤文本上训练的模型相当。此外,我们对内部表示的分析表明,语言模型通过将并列元素视为属于相似结构类别或通过类似删除的机制来处理不同类并列结构,这两者似乎都可以仅从接触同类并列结构中学习。这项工作有助于加深对语言模型如何在内部表示语言结构的理解,同时也通过展示模型在没有直接曝光的情况下如何泛化和处理不同类并列结构,为关于并列结构的更广泛辩论增添了内容。

英文摘要

Coordination, a fundamental linguistic structure, remains a subject of intense debate, and its exact nature continues to elude theoretical linguistics. A common view holds that only same-category constituents can be conjoined, which has been challenged by the many grammatical unlike coordinations found in natural language. Treating language models as a computational testbed, we investigate whether the acquisition of unlike coordination requires direct exposure in the training data, or whether it can emerge organically from general compositional abilities. Using Filtered-Corpus Training (FiCT), we train GPT-2 models on corpora from which all instances of unlike coordination have been removed. We find that direct exposure is not necessary: models trained on filtered data successfully generalize to unlike coordination, achieving perplexity and grammaticality judgments comparable to models trained on unfiltered text. Furthermore, our analyses of internal representations indicate that language models process unlike coordination by treating the conjoined elements as belonging to similar structural categories or through a mechanism akin to deletion, both of which appear learnable from exposure to alike coordination alone. This work contributes to the growing understanding of how language models internally represent linguistic structures, while also adding to the broader debate on coordination by showing how models generalize and process unlike coordination without direct exposure.

Comments13 pages, 6 tables, 2 figures, to submit to TACL

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑