为什么预训练无法共享跨语言知识
Why Pretraining Fails to Share Cross-Lingual Knowledge
- Weizmann Institute of Science(魏茨曼科学研究所)
- Bar-Ilan University(巴伊兰大学)
- Johns Hopkins University(约翰斯·霍普金斯大学)
- A*STAR(新加坡科技研究局)
- University of Washington(华盛顿大学)
- MIT(麻省理工学院)
- MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过受控双语预训练实验发现,不相交的token空间是跨语言知识泛化失败的根本原因,并提出通过逐词翻译映射到共享token空间,可恢复12.6%的母语学习效率。
AI中文摘要:
大型语言模型(LLMs)在处理和建模多种语言方面取得了显著进展。然而,与人类多语者不同,它们表现出令人惊讶的有限的跨语言知识迁移能力。虽然这一局限性已被充分记录,但其在多语言训练过程中的起源仍不清楚。我们预训练了360M和7B参数规模的LLMs,并表明较差的跨语言知识泛化在预训练期间出现,并在标准干预下持续存在。为了隔离其原因,我们采用了一种受控的双语预训练设置,使用同一语言的两种副本,共享相同的文本和分词,但映射到不相交的token空间。我们发现,即使在同一语言的相同副本之间,仅不相交的token就足以引发知识分隔,从而确立不相交的token空间是跨语言知识泛化的根本障碍。在这一理解的指导下,我们建议通过简单的逐词翻译将语言映射到共享的token空间,并发现这显著改善了跨语言知识泛化,恢复了母语学习效率的12.6%——是基线的14倍。
英文摘要:
Large Language Models (LLMs) have made remarkable progress in the processing and modeling of many languages. Yet, unlike human multilinguals, they exhibit surprisingly limited cross-lingual knowledge transfer. While this limitation is well documented, its origins during multilingual training remain unclear. We pretrain 360M- and 7B-parameter LLMs and show that poor cross-lingual knowledge generalization emerges during pretraining and persists under standard interventions. To isolate its cause, we employ a controlled bilingual pretraining setting using two copies of the same language, sharing identical text and token segmentation, but mapped to disjoint token spaces. We find that disjoint tokens alone are enough to induce knowledge compartmentalization, even between identical copies of the same language, establishing disjoint token spaces as a fundamental barrier to cross-lingual knowledge generalization. Guided by this understanding, we suggest mapping languages into a shared token space by simple word-wise translation and find it substantially improves cross-lingual knowledge generalization, recovering up to 12.6\% of native-language learning efficiency --- 14$\times$ the baseline.