建模完结性习得中的发展性转变
Modeling the Developmental Shift in Telicity Acquisition
- Michigan State University(密歇根州立大学)
- University of South Carolina(南卡罗来纳大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出惊异度差异方法自动标注完结性,发现儿童依赖动词后限定词线索而成人依赖动词类别,支持句法引导的习得轨迹。
AI中文摘要:
习得完结性——即有界事件(如“吃了一个苹果”)与无界事件(如“吃了苹果”)之间的区分——要求第一语言(L1)学习者将表层和语义线索映射到抽象事件结构上,但这一映射的计算轨迹尚不明确。我们提出了一种“惊异度差异”方法,利用GPT2在配对的时间状语诊断测试(“在一小时内”与“持续一小时”)上的token惊异度,自动标注英语CHILDES语料库中的完结性,并与专家语言学家的判断进行了验证。利用这些标签,我们在12个句法和词汇语义特征上训练了诊断性逻辑回归分类器,以比较儿童言语和儿童导向言语如何编码完结性。这两个模型出现分歧:儿童模型通过一个单一的确定性线索——动词后限定词的存在——达到了近乎完美的准确率,而成人模型则更依赖于动词类别和其他词汇语义特征,限定词线索被中和。这一轨迹支持句法引导(Syntactic Bootstrapping):学习者首先利用高频结构线索作为支架进行引导,然后才发展出完全组合性的、基于动词的事件结构。
英文摘要:
Acquiring telicity, which is the distinction between bounded (e.g., ate an apple) and unbounded (e.g., ate apples) events, requires first language (L1) learners to map surface-level and semantic cues to abstract event structures, but the computational trajectory of this mapping is not well understood. We introduce a Difference in Surprisal method that uses GPT2 token surprisal over paired temporal adverbial diagnostics (in an hour versus for an hour) to automatically label telicity across English CHILDES corpora, validated against expert linguist judgments. Using these labels, we train diagnostic logistic regression classifiers on 12 syntactic and lexical semantic features to compare how child speech and child-directed speech encode telicity. The two models diverge: the child model reaches near perfect accuracy through a single deterministic cue, the presence of a post-verbal determiner, while the adult model relies more heavily on verb class and other lexical semantic features, with the determiner cue neutralized. This trajectory supports Syntactic Bootstrapping: learners first exploit high-frequency structural cues as a scaffold to bootstrap, before developing fully compositional, verb-based event structures.