评估语言差异对大型语言模型多语言语法能力的影响
Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究通过四种评估方法在MultiBLiMP基准上测试六族模型,发现后期训练降低语法能力且低资源语言受损最重,母语提示可恢复其隐藏能力,故需多范式协议避免低估。
中文摘要 AI 辅助
关于多语言语言模型语法能力的论断,会因能力测量方式的不同而产生显著差异,然而,评估范式、后期训练与语言资源可用性之间的相互作用尚未得到系统性的检验。我们使用四种评估方法,在覆盖101种语言的句法最小对基准MultiBLiMP上,评估了来自六个模型家族的基础模型和后期训练模型。我们报告三项主要发现。第一,后期训练会降低语法能力,但这一影响的幅度因模型规模不同而不均匀地减小,而低资源语言承受的代价最高。第二,后期训练模型保留了无法通过显式提示表达出来的语法知识,但这仅在高级资源语言中可测量,因为低资源环境下的接近随机基线几乎没有可隐藏的知识。第三,母语提示能恢复低资源语言中原本隐藏的能力,这表明只有高资源语言才能直接从无提示的概率中进行探测。我们得出结论,多语言语法评估必须采用语言感知的多范式协议,以避免系统性地低估低资源语言的能力。
英文摘要
Claims about the grammatical competence of multilingual language models vary sharply with how competence is measured, yet the interaction between evaluation paradigm, post-training, and language resource availability has not been systematically examined. We evaluate base and post-trained models from six families on MultiBLiMP, a syntactic minimal-pair benchmark covering 101 languages, using four evaluation methods. We report three principal findings. First, post-training degrades grammatical competence, but the magnitude of this effect is reduced unevenly by model scale, while low-resource languages bear the highest cost. Second, post-trained models retain grammatical knowledge they cannot articulate through explicit prompting, yet this is measurable only in high-resource languages, because near-chance baselines in low-resource settings leave little knowledge to hide. Third, native-language prompting recovers otherwise hidden competence on low-resource languages, demonstrating that only high-resource languages can be probed directly from unprompted probabilities. We conclude that multilingual grammatical evaluation must adopt language-informed, multi-paradigm protocols to avoid systematically underestimating low-resource abilities.
发表机构
- University of Groningen(格罗宁根大学)
机构由 AI 辅助整理,请以论文原文为准。