M-GATE:面向大语言模型的多语言语法、翻译准确性与效率基准
M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models
AI总结:
本研究推出多语言基准M-GATE,评估50余种模型的80余种配置,发现翻译与语法能力分离,低资源语言惩罚随模型迭代缩小,推理对翻译提升显著。
AI中文摘要:
多语言语言模型被部署用于上百种甚至更多语言,但大多数基准测试仅检验模型是否能在某一语言中执行任务,而非检验模型对该语言本身的掌握程度,将流利度与熟练度混为一谈。我们推出M-GATE(多语言语法、翻译准确性与效率),这一基准覆盖30种类型多样的语言,从高资源语言到低资源语言不等,用于衡量语言熟练度。M-GATE包含三项任务:一是对语言学家精心设计、经对抗性选择的句子进行语法错误检测,这些句子聚焦于难以处理的、特定语言的现象;二是对共享英语源文本进行跨29种目标语言的往返翻译,由经专业标注员验证的三方LLM评审小组进行评分;三是补充的分词器效率度量。我们对50余种模型的80余种配置进行评估。流利度与熟练度存在明显差异:翻译能力出色的模型在对抗性语法项上的表现接近随机水平,最佳模型的马修斯相关系数(MCC)仅为0.36,且其错误倾向于系统性地未标记,接受不合语法的文本而非发出误报。翻译质量与预训练数据中某一语言的占比密切相关(与Common Crawl对数占比的相关系数r=0.86),产生了严重的低资源语言惩罚,但随着模型版本的迭代,这一差距正在缩小。启用推理能力可可靠地提升翻译质量,但其对错误检测的影响较小,且对部分模型为负向,因此最佳配置取决于任务类型。为防止测试数据被污染,测试项被隐藏在持续更新的公开排行榜之后,仅发布示例(此https URL)。
英文摘要:
Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency. We introduce M-GATE (Multilingual Grammar, Accuracy in Translation, and Efficiency), a benchmark of linguistic proficiency spanning 30 typologically diverse languages from high- to low-resource. M-GATE comprises three tasks: grammatical error detection on linguist-crafted, adversarially selected sentences that turn on hard, language-specific phenomena; round-trip translation of shared English sources across 29 target languages, scored by a three-provider LLM judge panel validated against professional annotators; and a supplementary tokenizer-efficiency measure. We evaluate over 50 models in more than 80 configurations. Fluency and proficiency come apart sharply: models that translate competently sit near chance on the adversarial grammar items, the best reaching a Matthews correlation coefficient (MCC) of only 0.36, and their errors lean systematically toward under-flagging, accepting ungrammatical text rather than raising false alarms. Translation quality closely tracks a language's share of pretraining data (r = 0.86 against log Common Crawl share), producing a steep low-resource penalty that is nonetheless narrowing with successive model releases. Enabling reasoning reliably improves translation, while its effect on error detection is smaller and for some models negative, so the best configuration is task-dependent. To resist contamination, test items are kept private behind a continuously updated public leaderboard, with illustrative examples released (https://m-gate.ai).