发表机构
ellamind; Leibniz University Hannover; University of Helsinki; University of Turku; ELLIS Institute Tübingen; Prior Labs; University of Oslo; Prompsit Language Engineering; AI Sweden; Instituto de Telecomunicações; Instituto Superior Técnico; TransPerfect; Charles University; Ontocord; LAION; Juelich Supercomputing Center (JSC), Research Center Juelich (FZJ)(Ellamind; 莱布尼茨汉诺威大学; 赫尔辛基大学; 图尔库大学; ELLIS 图宾根研究所; Prior Labs; 奥斯陆大学; Prompsit 语言工程公司; AI Sweden; 电信研究所; 高等技术学院; TransPerfect; 查理大学; Ontocord; LAION; 于利希超级计算中心 (JSC), 于利希研究中心 (FZJ))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MultiSynt/MT,一个覆盖36种欧洲语言、约4.8万亿token的合成平行语料库,通过翻译1000亿高质量token生成,使参考LLM在减少72%预训练token下达到原生数据基线性能,并发现标准基准测试的评估盲点。
AI 中文摘要
开放的网络规模预训练语料库仍然集中在英语上,限制了多语言LLM的发展。我们引入了MultiSynt/MT,一个开放的合成平行语料库,包含约4.8万亿目标语言token,覆盖36种欧洲语言,通过使用Tower+和OPUS-MT/HPLT-MT系统翻译1000亿高质量Nemotron-CC token生成。对于许多中低资源欧洲语言,这是最大的公开可用预训练资源。在一个广泛的多语言基准测试套件上,基于MultiSynt/MT训练的参考LLM达到了原生数据基线HPLT 2.0的最终分数,使用的预训练token减少了约72%,并在匹配的1000亿token训练预算下相对提升了约15%。我们的分析还识别出评估盲点:标准的多项选择基准测试无法捕捉翻译质量差异,而一种对流畅性敏感的LLM作为评判者的评估在训练的LLM上清晰地恢复了这些差异(MultiSynt本身没有流畅性缺陷),并且挪威语的习惯表达和文化基础任务仍然由原生数据更好地服务。我们发布了该语料库,包括来自多个系统的行对齐翻译,以支持对多语言预训练数据和评估的受控研究。
英文摘要
Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approximately 4.8 trillion target-language tokens across 36 languages, produced by translating 100 billion high-quality Nemotron-CC tokens with Tower+ and OPUS-MT/HPLT-MT systems. For many medium- and lower-resource European languages, this is the largest openly available pre-training resource. Across five high- and medium-resource languages, reference LLMs trained on MultiSynt/MT reach the final score of HPLT 2.0, a native-data baseline, using roughly 72% fewer pre-training tokens, and outperform it by approximately 15% relative at a matched 100B-token training budget. Our analyses also identify evaluation blind spots: standard multiple-choice benchmarks miss translation-quality differences that a fluency-sensitive LLM-as-judge protocol recovers on the trained LLMs without detecting a deficit relative to its native-data baseline, while Norwegian idiomatic and culturally grounded tasks remain better served by native data. We release the corpus, including row-aligned translations from multiple systems, to support controlled research on multilingual pre-training data and evaluation.
CommentsEMNLP 2026 Camera-ready Version