并行分词器:重新思考低资源语言跨语言转移中编码器模型的词汇设计
Parallel Tokenizers: Rethinking Vocabulary Design in Cross-Lingual Transfer of Low-Resource Languages
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对低资源语言跨语言转移中因词汇映射问题限制表示学习的情况,提出并行分词器框架,先单语训练再对齐词汇表,实验表明该方法训练的模型在多任务中优于传统基线,对推进多语言表示学习意义重大。
AI中文摘要:
分词是多语言模型的基础,但现有方法常因将语义等效词映射到不同嵌入而限制跨语言转移,如英语和豪萨语中的例子,在低资源语言中问题更突出。我们引入并行分词器框架,先单语训练分词器,再用双语词典或逐词翻译详尽对齐词汇表,这能强化跨语言共享语义空间并改善词频平衡。我们在13种低资源语言上从头预训练变压器编码器并评估,结果显示并行分词器训练的模型优于传统多语言基线,证实重新思考分词对推进多语言表示学习至关重要,尤其在低资源环境中。
英文摘要:
Tokenization forms the basis of multilingual language models, yet existing methods often limit cross-lingual transfer by mapping semantically equivalent words to different embeddings. For example, 'I eat rice' in English and 'Ina cin shinkafa' in Hausa are typically mapped to different vocabulary indices, preventing shared representations and limiting cross-lingual generalization. This problem is even more pronounced in low-resource languages, where shared representations could offer the greatest benefit. We introduce parallel tokenizers, a new framework that first trains tokenizers monolingually and then aligns their vocabularies exhaustively using bilingual dictionaries or word-to-word translation. This alignment enforces a shared semantic space across languages while naturally improving fertility balance. To assess their effectiveness, we pretrain a transformer encoder from scratch on thirteen low-resource languages and evaluate it on sentiment analysis, hate speech detection, emotion classification, and sentence embedding similarity. Across all tasks, models trained with parallel tokenizers outperform conventional multilingual baselines, confirming that rethinking tokenization is essential for advancing multilingual representation learning--especially in low-resource settings.