arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25904cs.CL

一种实现跨语言迁移的统一形式:超越原生正字法的多语言语言模型预训练

One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography

  • Ohio State University(俄亥俄州立大学)
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

Muge Zhang, Aaron Jencks, Krishna Badikela, Yulia Tsvetkov, Sachin Kumar

AI总结:

本文通过系统对比三种输入表示,发现罗马化预训练可实现最强跨语言迁移,且优势随模型规模增大而扩大,提出多语言模型应将罗马化作为预训练核心设计选择的结论。

AI中文摘要:

多语言语言模型通过共享子词词汇在不同语言间迁移知识,当相关语言使用不同书写系统时,该机制会失效。现有研究通过脚本均衡化(罗马化或国际音标转录)解决这一问题,但直接对比研究较少,且多聚焦于仅编码器模型,多数工作仅适配现有预训练模型。本文在自回归多语言预训练中系统对比不同输入表示,在受控设置下针对四个类型学配对中的八种语言,对比正字法文本、国际音标(IPA)和罗马化三种形式,覆盖三个模型规模(4.67亿、7.09亿和10.3亿参数)。在可见和不可见语言的广泛下游任务中,罗马化预训练实现了最强的跨语言迁移效果,且该优势随模型规模增大而扩大;国际音标在多数场景下优于正字法文本,但落后于罗马化。令人意外的是,对文本预训练模型在罗马化数据上进行微调,会损害基础模型已覆盖语言的性能,仅在模型缺乏对应脚本覆盖时才略微有帮助。研究结果表明,对于涵盖类型学多样脚本的多语言模型,要获得最大收益,应将罗马化视为预训练阶段应用的核心设计选择,而非事后补救措施。

英文摘要:

Multilingual language models transfer knowledge across languages through shared subword vocabulary, a mechanism that breaks down when related languages use different writing systems. Prior work addresses this via script equalization (romanization or IPA transcription), but direct comparisons are rare; the focus has been on encoder-only models, with most work adapting existing pretrained models. We systematically compare different input representations in autoregressive multilingual pretraining, comparing orthographic text, IPA, and romanization in a controlled setup across three scales (467M, 709M, and 1.03B) on eight languages in four typologically motivated pairs. Across a wide range of downstream tasks on seen and unseen languages, romanized pretraining yields the strongest cross-lingual transfer, and the advantage over text widens with scale. IPA improves over text in most settings but trails romanization. Surprisingly, finetuning a text-pretrained model on romanized data hurts performance on languages already covered by the base model, only marginally helping when the model lacks script coverage. Our results indicate that for multilingual models spanning typologically diverse scripts, to obtain maximum benefits, romanization should be treated as a core design choice applied at pretraining rather than a post hoc fix.

补充信息

↑