语言平等是有代价的:针对欧盟24+语言的多轮大语言模型性能系统研究
Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+
AI总结:
本研究以30种语言评估9种LLM,发现开源模型难以覆盖欧盟24种官方语言,商用模型表现更优且非英语语言运行成本更高、得分更低,揭示语言平等存在代价。
AI中文摘要:
我们将大语言模型(LLM)作为语言智能体,在30种语言(欧盟24种官方语言外加6种其他语言)中进行自博弈的目标导向对话游戏评估。与静态或基于偏好的评估不同,该范式是多轮、无参考且可程序化评分的,且由于游戏机制与语言无关,只需本地化一组固定的提示词和词表文件即可扩展至新语言。我们评估了9种开源权重和商用LLM,发现没有开源权重模型能很好覆盖欧盟24种语言:在每种官方语言中,两款商用系统的得分均超过所有开源权重模型,且两款得分最低的模型在欧盟24种语言中的平均得分低于40分。即使在公开网络文本量少四个数量级的语言中,商用系统仍保持领先,表明语言对等是可实现的,但仅靠公开爬虫数据无法达成。模型的所属地区会提升其性能,但无法缩小差距:对于两款中国开发的模型,中文是30种语言中表现最强的,但所有模型中最佳的中文得分属于一款美国商用系统。覆盖范围也并非服务对等:在模型和语言的汇总统计中,非英语语言的运行成本中位数比英语高31%,得分则低10%。
英文摘要:
We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU languages plus six others. Unlike static or preference-based evaluation, this paradigm is multi-turn, reference-free and programmatically scored, and because the game mechanics are language-agnostic it extends to a new language by localising a fixed set of prompt and word-list files. Evaluating nine open-weight and commercial LLMs, we find that no open-weight model covers the EU-24 well: in every official language both commercial systems outscore every open-weight model, and the two weakest average below 40 points across the EU-24. The commercial systems stay ahead even in languages with four orders of magnitude less public web text, showing that linguistic parity is achievable, but not from public crawls alone. A model's home region lifts it without closing the gap: Chinese is the strongest of all 30 languages for two Chinese-developed models, yet the best Chinese score of any model belongs to a US commercial system. Coverage is also not parity of service. Pooled over models and languages, the median non-English language costs 31% more to run than English, and scores 10% lower.