适配粤语的语言模型能否更好地预测粤语阅读?一项跨模型眼动评估
Do Cantonese-Adapted Language Models Better Predict Cantonese Reading? A Cross-Model Eye-Tracking Evaluation
- The Hong Kong Polytechnic University(香港理工大学)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过自然粤语眼动数据,对比两组同系列适配模型,发现更广泛的粤语专属训练模型CantoneseLLM-7B的预测拟合度更强,模型排序取决于所用信息论指标。
AI中文摘要:
源自自回归语言模型的信息论指标被广泛用于刻画塑造人类阅读的预期,但针对特定语言变体的训练是否会提升这种心理语言学对齐效果仍不明确。对于粤语而言,该问题尚无定论:近期NLP评估显示,相较于面向普通话的模型或通用模型,粤语专属训练带来的收益参差不齐。本研究利用自然粤语眼动数据,对比两组同系列适配模型:CKIP GPT-2 Tiny与其经轻度粤语适配的衍生模型JED351,以及Qwen2.5-7B与经更广泛粤语持续预训练和指令调优的CantoneseLLM-7B。从各模型中提取词汇意外性、词性意外性、目标前熵及熵减少量。词汇意外性与包含四项指标的联合模型始终将排序偏向CantoneseLLM-7B,其次为Qwen2.5-7B、CKIP、JED351,而熵减少量则偏向CKIP。这些结果表明,更广泛的粤语专属训练或与更强的预测拟合度相关,同时模型排序也取决于所评估的信息论指标。
英文摘要:
Information-theoretic measures derived from autoregressive language models are widely used to characterize the expectations that shape human reading, but whether language-variety-specific training improves such psycholinguistic alignment remains unclear. This question is still open for Cantonese, where recent NLP evaluations reported mixed benefits from Cantonese-specific training relative to Mandarin-oriented or general-purpose models. Using naturalistic Cantonese eye-tracking data, we compare two within-family adaptation contrasts: CKIP GPT-2 Tiny versus its lightly Cantonese-adapted JED351 derivative, and Qwen2.5-7B versus CantoneseLLM-7B, which underwent substantially more extensive Cantonese continued pretraining and instruction tuning. From each model, we derive lexical surprisal, POS surprisal, entropy before the target, and entropy reduction. Lexical surprisal and the joint four-metric model consistently favor CantoneseLLM-7B, followed by Qwen2.5-7B, CKIP, and JED351, whereas entropy reduction favors CKIP. These results suggest that more extensive Cantonese-specific training can be associated with stronger predictive fit, while model rankings also depend on the information-theoretic measure being evaluated.