arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

句法结构的无数据通用先验

A Data-free Universal Prior over Syntactic Structures

Fermín Moscoso del Prado Martín

arXiv 2609.16854首次发表:更新:

发表机构

University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于认知的增量语言产出模型,无需语言数据即可产生句法结构的通用先验,在138种语言中优于随机树,并与语料库概率正相关,揭示句法概率的认知起源。

AI 中文摘要

概率是语言理解、产出、习得和演化理论以及大型语言模型的基础。现有理论从特定语言的数据中估计句法结构的概率。这种概率结构的一部分是否能在不依赖特定语言经验的情况下产生,仍然未知。本文表明,一种基于认知的增量语言产出模型——其中词汇通过网络增长逐步整合到句法结构中——会产生一种句法结构的通用先验。所得先验在不将参数拟合到语言数据的情况下,为句法结构(表示为依存树)分配概率,并且在所有138种类型多样的语言中,为已证实的树分配的概率高于随机树。在34种语言中的33种中,这些先验概率与从语料库中估计的概率呈正相关。结果表明,句法概率结构的一部分可以独立于特定语言的统计学习而产生。因此,语言经验可能优化那些已经由语言产出过程结构化的概率,而不是从一个最初均匀的空间中创造它们。这为句法结构概率分布的一部分确定了可能的认知起源,将语言产出与统计学习联系起来,同时为概率语言模型提供了一种数据无关的结构性偏差。

英文摘要

The probabilities of syntactic structures in human languages are assumed to emerge fully from language-specific experience. Here, I show that a universal prior over syntactic structures emerges from a model of human language production, in which words are progressively integrated into syntactic structure. Without fitting any parameters to specific language data, the resulting prior assigns higher probabilities to attested than to random dependency trees in all 138 typologically diverse languages examined. These prior probabilities correlate positively with those estimated from corpora in 33 of 34 languages. The results indicate that part of the probability structure of syntax can arise independently of language-specific learning. This identifies human language production as a possible cognitive source of universal statistical structure in language, while providing a data-independent structural bias for probabilistic models, including large language models.

Comments30 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑