语言模型惊异度对中文阅读预测能力的系统分析
A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese
- Shanghai Jiao Tong University(上海交通大学)
- Nanyang Technological University(南洋理工大学)
- Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究系统分析语言模型惊异度对中文阅读时间的预测能力,提出SMS对齐方案,发现预测效果因语料库而异,并揭示逆缩放现象及n-gram接近性解释。
AI中文摘要:
本研究分析了基于语言模型的词级惊异度对普通话阅读时间的预测能力。我们首先提出了最短匹配序列(SMS),一种对齐方案,用于在眼动追踪语料库所假设的词分割与语言模型的子词分词之间进行映射,因为在普通话背景下这两种分词方式常常不一致。然后,使用一系列从头训练、包含300亿词元的Chinese-Pythia模型(1400万至14亿参数),我们检验了惊异度在三个段落级普通话眼动追踪语料库(GECO-CN、HKP和MECO)中对首次注视持续时间、凝视持续时间和总阅读时间的预测效果。与先前零结果相反,我们的结果表明惊异度能够预测中文阅读时间。然而,预测能力是否随模型大小和训练量扩展则因语料库而异:在GECO-CN中,更大的模型预测更好,而在HKP以及MECO的最大规模中出现了逆缩放现象。随后,我们测试了HKP中逆缩放的一种可能解释,发现惊异度更接近n-gram统计量的检查点对阅读的预测更好。总而言之,惊异度对中文阅读时间测量的预测能力具有语料库特异性,这提醒我们不应仅从单一语料库得出缩放结论。
英文摘要:
This study analyzes the predictive power of LM-derived, token-level surprisal on Mandarin Chinese reading times. We first propose the Shortest Matching Sequence (SMS), an alignment scheme that maps between the word segmentation assumed by eye-tracking corpora and the LMs' subword tokenization, as the two tokenizations often disagree in the context of Mandarin Chinese. Then, using a suite of Chinese-Pythia models (14M-1.4B) trained on scratch with 30B tokens, we examine how well surprisal predicts first fixation duration, gaze duration, and total reading time in three paragraph-level eye-tracking corpora of Mandarin Chinese (GECO-CN, HKP, and MECO). Contrary to previous null findings, our results show that surprisal is predictive of Chinese reading times. However, whether predictive power scales with model size and the amount of training is corpus-specific: bigger models predict better in GECO-CN, whereas inverse scaling emerges in HKP and, at the largest sizes, in MECO. Subsequently, we tested one possible explanation for the inverse scaling in HKP and found that checkpoints whose surprisal remains closer to $n$-gram statistics are better predictors of reading. All in all, the predictive power of surprisal on Chinese reading time measurements is corpus-specific, which cautions against drawing scaling conclusions from a single corpus.