发表机构
ETH Zürich; Johns Hopkins University(苏黎世联邦理工学院; 约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过理论证明与小规模经验研究,指出嵌入空间结构不存在多语言性的理论诅咒,多语言模型性能下降的经验诅咒源于现实数据与训练条件。
AI 中文摘要
多语言自然语言处理(NLP)的核心目标是用一个多语言模型实现每种语言的高单语性能,同时具备跨语言对齐能力,以覆盖大规模语言。多语言性诅咒描述的是随着语言覆盖范围扩大,多语言模型性能下降的现象,对上述目标构成威胁。本文探究多语言嵌入空间是否天生无法在不大幅增加所需容量的情况下实现完美多语言性。我们首先将“完美多语言性”的目标形式化为两个多语言性条件,随后证明实现完美多语言性所需的最小维度仅随语言数量呈对数增长,即嵌入空间结构不存在多语言性的理论诅咒。这表明多语言性的经验诅咒是现实世界数据和训练条件导致的结果。我们通过一项小规模经验研究佐证了这一理解。本文首次从理论和内在视角研究多语言性诅咒,对该现象的科学理解具有重要意义。
英文摘要
A central goal of multilingual NLP is to achieve high monolingual performance per language and cross-lingual alignment for large-scale language coverage with a multilingual model. The curse of multilinguality describes the phenomenon of degradation in multilingual model performance as we increase language coverage, posing a threat to the above goal. This paper asks whether multilingual embedding spaces are inherently incapable of achieving perfect multilinguality without a prohibitive increase in required capacity. We first formalize the goal of "perfect multilinguality", embodied in two multilinguality conditions. We then prove that the minimum dimensionality required for perfect multilinguality grows only logarithmically in the number of languages. That is, we show that there is no theoretical curse of multilinguality for embedding space structure. This suggests that the empirical curse of multilinguality is a result of real world data and training conditions. We back this understanding with a small-scale empirical study. Our paper provides the first theoretical and intrinsic perspective on the curse of multilinguality.