arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从暴露到预期:西班牙语发展中的频率、意外性与语言

From Exposure to Expectation: Frequency, Surprisal, and Language Across Development in Spanish

Francisco Portillo López

arXiv 2608.22452首次发表:更新:

发表机构

Universidad de Navarra(纳瓦拉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过两项西班牙语语料库研究,发现频率可预测儿童词汇习得年龄,意外性可预测成人阅读加工难度,二者在语言发展中作用不同。

AI 中文摘要

意外性(Surprisal,语言模型给定前文语境分配给某个词的负对数概率)可可靠预测成人阅读时长。它在解释儿童习得单个词汇的时间方面是否也有同等作用?频率反映学习者对某个词的累积暴露,而意外性则反映该词在特定语境下单次出现的可预测性。我们通过两项基于语料库的西班牙语研究探究该问题。研究1中,我们使用儿童导向言语中的词汇频率、语境多样性,以及三种架构和训练语言不同的语言模型(BETO、BERTIN、mGPT)的意外性,对225个西班牙语名词的习得年龄(AoA)进行建模。频率对AoA有强预测作用(r=-.597,p<.001);除频率和词长外,意外性的贡献很小,在自然语境分析中亦是如此。研究2中,我们使用多语言眼动语料库(MECO Wave 2)的智利西班牙语子样本中的成人注视时长,结合mGPT意外性与两种独立频率测量指标进行建模。在控制频率和词长后,意外性可稳健预测更长的注视时长,且在两种频率来源间一致。匹配的词型水平比较显示,意外性与行为的关联在阅读中比在语言习得中更强(z=3.63,p<.001)。研究结果表明,累积词汇暴露与语境可预测性在语言发展轨迹中发挥不同作用:频率对早期词汇表征的习得时间尤其具有预测性,而意外性则捕捉已建立的语言系统中逐时刻的加工难度。我们结合基于使用的词汇发展理论、基于固化的词汇发展理论,以及语言模型作为人类语言行为模型的评估,对该模式进行讨论。

英文摘要

Surprisal, the negative log-probability a language model assigns to a word given its preceding context, reliably predicts adult reading times. Does it contribute as much to explaining when children acquire individual words? Frequency reflects a learner's cumulative exposure to a word, whereas surprisal reflects how predictable a single occurrence is given its context. We investigate this question across two corpus-based studies of Spanish. In Study 1, we modeled age of acquisition (AoA) for 225 Spanish nouns using lexical frequency and contextual diversity from child-directed speech, plus surprisal from three language models differing in architecture and training language (BETO, BERTIN, mGPT). Frequency strongly predicted AoA (r=-.597, p<.001); surprisal added little beyond frequency and word length, including in a naturalistic-context analysis. In Study 2, we modeled adult fixation durations in the Chilean Spanish subsample of the Multilingual Eye-movement Corpus (MECO Wave 2), using mGPT surprisal alongside two independent frequency measures. Surprisal robustly predicted longer fixation durations after controlling for frequency and word length, consistent across both frequency sources. A matched word-type-level comparison showed the surprisal-behavior association was stronger in reading than in acquisition (z=3.63, p<.001). The findings suggest cumulative lexical exposure and contextual predictability play different roles across the language trajectory: frequency is particularly informative about when early lexical representations are acquired, whereas surprisal captures moment-to-moment processing difficulty in an already-established linguistic system. We discuss this pattern in relation to usage-based and entrenchment-based accounts of lexical development and to the evaluation of language models as models of human language behavior.

Comments30 pages; 4 figures; 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑