发表机构
Ruzivo Research Lab; African Leadership University(鲁齐沃研究实验室; 非洲领导力大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对NLLB-200等多语言NMT模型未见低资源语言的代理token选择问题,提出多语言词嵌入初始化策略,在林布语-英语翻译上取得与最佳单语言代理相当的性能,消除了启发式选择需求,但存在声调变音符号保留的挑战。
AI 中文摘要
NLLB-200等多语言神经机器翻译模型覆盖200种语言,但仍有数千种语言未被支持,包括喀麦隆的大多数草原班图语。当为某种未见语言微调这些模型时,从业者必须选择一个代理语言token,但目前尚无原则性方法用于此选择。我们实现了一种词嵌入初始化策略,其中语言token是模型中已有的多个类型学相关语言的嵌入的平均值。我们使用来自新约文本的8837个句子对的平行语料库和双语词典,在林布语-英语翻译上评估该方法。我们比较了以下模型:NLLB-200零样本模型(chrF2++=12.5)、从头训练的Transformer模型(chrF2++=14.5)、用斯瓦希里语代理token微调的NLLB-200(chrF2++=47.3),以及采用我们的平均词嵌入初始化的NLLB-200(chrF2++=46.7)。我们发现多语言初始化的性能与最佳单语言代理相当,两种NLLB-200变体均比从头训练的基线模型提升了超过32个chrF2++点。这些结果表明,在极低资源的班图语翻译中,多语言迁移是主导因素,同时消除了对启发式代理选择的需求。然而,所有系统都无法保留声调变音符号,这凸显了一个开放的挑战。我们公开了我们的数据集和代码,以支持进一步的研究。
英文摘要
Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon. When fine-tuning these models for an unseen language, practitioners must choose a proxy language token, yet no principled method exists for this selection. We implemented an embedding initialization strategy where a language token is the average of embeddings from multiple typologically related languages already in the mod el. We evaluate this approach on Limbum-to-English translation using a parallel corpus of 8,837 sentence pairs from New Testament text and a bilingual dictionary. We compare models: NLLB-200 zero-shot (chrF2++ = 12.5), a Transformer trained from scratch (chrF2++ = 14.5), NLLB-200 fine-tuned with a Swahili proxy token (chrF2++ = 47.3), and NLLB-200 with our averaged embedding initialization (chrF2++ = 46.7). We find that the multi-language initialization achieves performance comparable to the best single-language proxy. Both NLLB-200 variants improve over the from-scratch baseline by over 32 chrF2++ points. These results show that multilingual transfer is the dominant factor in extremely low-resource Bantu translation while eliminating the need for heuristic proxy selection. However, all systems fail to preserve tonal diacritics, highlighting an open challenge. We make our dataset and code available to support further research.