公平可比?面向可比较的跨语言语言模型评估
Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation
- Georgetown University(乔治城大学)
- EleutherAI
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对跨语言语言模型评估的可比性挑战,通过受控单语言模型和多语言LLM验证,发现常用归一化指标存在跨语言偏差,提出语义等价序列的句子级负对数似然更适合跨语言比较。
AI中文摘要:
跨语言语言模型的公平可比评估仍是多语言自然语言处理的核心挑战。现有研究采用多种下游任务和内在指标,其理论依据各不相同,但鲜有实证研究考察这些方法是否能得出有意义的跨语言结论。我们系统研究了跨语言评估方法,使用在平行数据上训练的受控单语言语言模型,这些模型具有不同的分词器词汇量和模型规模,并在多语言大型语言模型(LLM)上验证了我们的发现。我们进一步探讨了实现跨语言可比下游评估面临的挑战。我们的结果表明,几种广泛使用的归一化指标会引入源于分词、编码和正字法差异的跨语言偏差。相比之下,基于语义等价序列计算的句子级负对数似然能提供更有意义且一致的跨语言比较。
英文摘要:
Crosslingual evaluation of language models that enables fair comparisons remains a fundamental challenge in multilingual NLP. Existing studies adopt a variety of downstream tasks and intrinsic metrics with different theoretical justifications, yet there has been little empirical investigation into whether these approaches yield meaningful crosslingual conclusions. We systematically examine crosslingual evaluation approaches using controlled monolingual language models trained on parallel data with varying tokenizer vocabulary sizes and model sizes, and further validate our findings on multilingual LLMs. We further discuss challenges in achieving comparable downstream evaluation across languages. Our results show that several widely used normalized metrics introduce crosslinguistic biases rooted in tokenization, encoding, and orthographic differences. In contrast, sentence-level negative log-likelihood computed over semantically equivalent sequences provides more meaningful and consistent crosslingual comparisons.