arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩展基于吉兹语的非洲语言词汇:阿姆哈拉语和提格雷尼亚语的比较研究

Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya

Hailay Kidu Teklehaymanot, Debela Desalegn Yadeta, Wolfgang Nejdl

arXiv 2607.15209首次发表:更新:

AI 中文总结

针对低资源非拉丁脚本语言表现不佳问题,提出VEXMLM。通过训练特定语言分词器扩展XLM - R词汇并初始化嵌入,分两阶段训练,在多种任务上评估。贡献为定制程序、两阶段策略及多语言评估,提升了相关语言任务性能。

AI 中文摘要

多语言预训练语言模型(PLMs)在低资源、非拉丁脚本语言上表现不佳,因为以拉丁脚本为中心的分词器训练导致词汇外(OOV)率高和子词过度碎片化。我们引入了VEXMLM,它是XLM - R的词汇扩展变体,针对资源最丰富的两种吉兹语脚本语言阿姆哈拉语和提格雷尼亚语,并在另外17种低资源非洲语言上进行了评估。我们在精心策划的阿姆哈拉语和提格雷尼亚语单语语料库上训练特定语言的SentencePiece分词器,用从该分词器派生的30000个吉兹语脚本子词扩展XLM - R的词汇,并通过平均XLM - R原始分词器下其组成子词的嵌入来初始化它们的嵌入。VEXMLM分两个阶段训练:(1)在精心策划的语料库上对扩展词汇进行持续的掩码语言建模,(2)在问答(QA)、命名实体识别(NER)和情感分析(SA)上进行监督微调。在阿姆哈拉语/提格雷尼亚语QA上,VEXMLM达到87.0 EM / 90.0 F1,而XLM - R为66.0 EM / 78.0 F1,Glot500为74.0 EM / 78.0 F1。在SA上,VEXMLM达到80.0%的准确率,而XLM - R为77.0%,Glot500为46.0%。在NER上,对于19种评估语言中的11种进行OOV分析时,VEXMLM将OOV令牌实体准确率从81.4%提高到94.3%。我们的贡献包括:(i)针对吉兹语脚本量身定制的词汇扩展和嵌入初始化程序;(ii)两阶段训练策略,阿姆哈拉语/提格雷尼亚语的词汇和持续预训练收益转移到17种类型相关的未增强非洲语言;(iii)跨越所有19种语言的内在分词指标(词汇覆盖率、生育率、OOV率)和外在任务性能的评估。

英文摘要

Multilingual pre-trained language models such as XLM-R perform well for major languages but struggle with low-resource Ge'ez-script languages, largely because Latin-script-centric tokenizers split their words into many subwords. We introduce VEXMLM, a vocabulary-extended variant of XLM-R targeting Amharic and Tigrinya. We train language-specific SentencePiece tokenizers on monolingual corpora, extend XLM-R's vocabulary with 30k Ge'ez-script subwords, and initialize each new embedding to the mean of the pretrained embeddings. VEXMLM undergoes two-stage training: (1) continued masked language modeling on the monolingual corpora and (2) supervised fine-tuning on question answering and named entity recognition (Amharic and Tigrinya) and sentiment analysis (Amharic). VEXMLM lowers tokenizer fertility below that of XLM-R and Glot500 on both languages, by 28.0% (Amharic) and 45.9% (Tigrinya) relative to XLM-R. Downstream, it modestly improves named entity recognition over XLM-R, scores below XLM-R on extractive question answering, and is comparable on sentiment analysis. An ablation on Tigrinya NER shows that vocabulary expansion alone lowers accuracy on out-of-vocabulary words (words that XLM-R's tokenizer cannot represent or splits into more pieces than the expanded tokenizer), and that continued pretraining is required for the expanded model to exceed the baseline. Vocabulary expansion thus makes Ge'ez-script tokenization substantially more efficient, while its downstream benefit depends on the task and on adapting the new embeddings through continued pretraining. Resources: GitHub repository | Hugging Face model.

Comments12 pages , 5 tables , 1 figurs

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑