arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MultiSynt/MT:跨36种语言翻译的万亿Token多平行预训练数据

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

Maximilian Idahl, Jörg Tiedemann, Sampo Pyysalo, David Salinas, Tomasz Galica, Shenbin Qian, Tudor Nicolae Mateiu, Zihao Li, Anna Lokrantz, Fedor Vitiugin, André F. T. Martins, Jenna Kanerva, Filip Ginter, Matthias Lindemann, Tim Isbister, Birger Moell, Jonas Lindh, Jan Hajič, Jenia Jitsev, Andrey Kutuzov, Stephan Oepen, Gema Ramírez-Sánchez

arXiv 2607.00890首次发表:更新:

发表机构

ellamind; Leibniz University Hannover; University of Helsinki; University of Turku; ELLIS Institute Tübingen; Prior Labs; University of Oslo; Prompsit Language Engineering; AI Sweden; Instituto de Telecomunicações; Instituto Superior Técnico; TransPerfect; Charles University; Ontocord; LAION; Juelich Supercomputing Center (JSC), Research Center Juelich (FZJ)(Ellamind; 莱布尼茨汉诺威大学; 赫尔辛基大学; 图尔库大学; ELLIS 图宾根研究所; Prior Labs; 奥斯陆大学; Prompsit 语言工程公司; AI Sweden; 电信研究所; 高等技术学院; TransPerfect; 查理大学; Ontocord; LAION; 于利希超级计算中心 (JSC), 于利希研究中心 (FZJ))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MultiSynt/MT,一个覆盖36种欧洲语言、约4.8万亿token的合成平行语料库,通过翻译1000亿高质量token生成,使参考LLM在减少72%预训练token下达到原生数据基线性能,并发现标准基准测试的评估盲点。

AI 中文摘要

开放的网络规模预训练语料库仍然集中在英语上,限制了多语言LLM的发展。我们引入了MultiSynt/MT,一个开放的合成平行语料库,包含约4.8万亿目标语言token,覆盖36种欧洲语言,通过使用Tower+和OPUS-MT/HPLT-MT系统翻译1000亿高质量Nemotron-CC token生成。对于许多中低资源欧洲语言,这是最大的公开可用预训练资源。在一个广泛的多语言基准测试套件上,基于MultiSynt/MT训练的参考LLM达到了原生数据基线HPLT 2.0的最终分数,使用的预训练token减少了约72%,并在匹配的1000亿token训练预算下相对提升了约15%。我们的分析还识别出评估盲点:标准的多项选择基准测试无法捕捉翻译质量差异,而一种对流畅性敏感的LLM作为评判者的评估在训练的LLM上清晰地恢复了这些差异(MultiSynt本身没有流畅性缺陷),并且挪威语的习惯表达和文化基础任务仍然由原生数据更好地服务。我们发布了该语料库,包括来自多个系统的行对齐翻译,以支持对多语言预训练数据和评估的受控研究。

英文摘要

Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approximately 4.8 trillion target-language tokens across 36 languages, produced by translating 100 billion high-quality Nemotron-CC tokens with Tower+ and OPUS-MT/HPLT-MT systems. For many medium- and lower-resource European languages, this is the largest openly available pre-training resource. Across five high- and medium-resource languages, reference LLMs trained on MultiSynt/MT reach the final score of HPLT 2.0, a native-data baseline, using roughly 72% fewer pre-training tokens, and outperform it by approximately 15% relative at a matched 100B-token training budget. Our analyses also identify evaluation blind spots: standard multiple-choice benchmarks miss translation-quality differences that a fluency-sensitive LLM-as-judge protocol recovers on the trained LLMs without detecting a deficit relative to its native-data baseline, while Norwegian idiomatic and culturally grounded tasks remain better served by native data. We release the corpus, including row-aligned translations from multiple systems, to support controlled research on multilingual pre-training data and evaluation.

CommentsEMNLP 2026 Camera-ready Version

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑