arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34936cs.CL

神经语言模型学习依存结构的上下文分布:关于组合性的统计学习理论

Neural Language Models Learn the Contextual Distributions of Dependency Structures: a statistical learning theory to compositionality

Wang Bojun, Junjie Chen, Holly Jenkins, Elizabeth Wonnacott

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出统计学习理论,认为NLM将已学依存结构作为新分布单元,通过合成语法实验证明模型能学习复合结构的上下文分布,且依存关系学习先于上下文特征,为组合性提供解释。

中文摘要 AI 辅助

目前尚不清楚神经语言模型(NLMs)如何获取由语法结构编码的、独立于词汇语义的结构意义。我们提出一个统计学习过程,在该过程中,已学习的依存结构本身成为后续统计学习的新的分布单元。在这种解释下,一旦获得依存结构,模型就会跟踪其上下文分布。这些上下文特征反映了复合结构的语义属性。为验证这一假设,我们设计了一种合成语法,其中每个语法结构具有不同的上下文分布,这些分布无法仅从组成词元的分布统计中恢复。我们在此语法上训练一系列BERT风格的掩码语言模型,并考察其发展轨迹。结果表明,即使无法仅从词元统计中推断出复合依存结构的上下文分布,模型也能成功学习它们。发展分析进一步揭示了清晰的发展轨迹:定义语法结构的依存关系的学习始终先于其上下文特征的学习。这些发现表明,NLM中的统计学习不仅仅是词元共现统计的累积,而是一个已学习的依存结构成为分布学习新单元的过程。我们认为,这一过程为NLM如何解决语言中的组合性问题提供了一种统计学习的解释。最后,我们讨论了这一统计学习过程可能为语言认知如何从纯粹分布统计中涌现提供解释性理论的可能性。

英文摘要

It is unclear how Neural Language Models (NLMs) acquire the structural meaning encoded by grammatical structures that is independent of lexical semantics. We propose a statistical learning process in which learned dependency structures themselves become new distributional units for subsequent statistical learning. Under this account, once a dependency structure is acquired, the model tracks its contextual distributions. These contextual features reflect the semantic properties of a composite structure. To test this hypothesis, we design a synthetic grammar in which each grammatical structure has distinct contextual distributions that cannot be recovered from the distributional statistics of their component tokens alone. We train a series of BERT-style masked language models on this grammar and examine their developmental trajectory. The results show that models can successfully learn the contextual distributions of composite dependency structures even though they cannot be inferred from token statistics alone. Developmental analysis further reveals a clear developmental trajectory. The learning of the dependency relations that define a grammatical structure consistently precedes the learning of its contextual features. These findings suggest that statistical learning in NLMs is not merely the accumulation of token co-occurrence statistics, but a process in which learned dependency structures become new units of distributional learning. We argue that this process provides a statistical-learning account of how NLMs solve the compositionality problem in language. Finally, we discuss the possibility that this statistical learning process provides an explanatory theory on how language cognition could emerge from pure distributional statistics.

发表机构

  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑