AI 中文总结
该研究对比儿童与语言模型的词汇学习,发现儿童词汇学习呈加速积累特征,而语言模型则呈现恒定比例回报,且儿童学习所需训练数据远少于语言模型。
AI 中文摘要
儿童在生命最初几年里学习数百个单词,这一过程起初进展缓慢,但很快会加速。先前的模型将词汇增长描述为随时间推移的证据积累。本文表明,该过程最适合被描述为加速积累:儿童从每增加一个单位的语言经验中获得的知识,多于从前一个单位经验中获得的。与儿童不同,语言模型——即使是在面向儿童的语音上训练的模型——也不会表现出加速,相反,它们对新数据呈现出恒定的比例回报,这与缩放定律一致。儿童学习所使用的训练数据比语言模型少几个数量级;他们对学习输入日益高效的利用是一个有待研究的候选解释。
英文摘要
Children learn hundreds of words over the first years of their lives, in a process that begins slowly but quickly picks up speed. Prior models describe vocabulary growth as evidence accumulation over time. Here we show that the process is best characterized as accelerating accumulation: children learn more from each additional unit of linguistic experience than they did from the one before. In contrast to children, language models -- even those trained on child-directed speech -- do not accelerate. Instead, they show constant proportional returns on new data, consistent with scaling laws. Children learn using many orders of magnitude less training data than language models; their increasingly efficient use of their learning input is a candidate explanation.