arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KinyaEmbed:基于多阶段课程训练的卢旺达语(Kinyarwanda)对比句嵌入模型

KinyaEmbed: Contrastive Sentence Embeddings for Kinyarwanda via Multi-Stage Curriculum Training

Ireddi Rakshitha, Devavarapu Yashwanth, Ntakirutimana Pierre

arXiv 2608.26941首次发表:更新:

发表机构

Barclays; Carnegie Mellon University(巴克莱银行; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对卢旺达语预训练嵌入不足的问题,提出基于KinyaBERT-large的KinyaEmbed模型,经四阶段课程训练后,在SemRel2024-rw等基准上超越现有多语言模型,效果显著且相关资源公开。

AI 中文摘要

我们提出了KinyaEmbed,这是首个专门针对卢旺达语(Kinyarwanda)的句嵌入模型,卢旺达语是一种形态丰富的班图语,在卢旺达有超过1200万人使用。现有的多语言嵌入模型如LaBSE、mE5-large和OpenAI text-embedding-3-large在卢旺达语上表现不佳,因为它们的预训练语料库中该语言的代表性严重不足。KinyaEmbed基于KinyaBERT-large构建,通过使用MultipleNegativesRankingLoss(MNRL)的四阶段课程进行训练:第一阶段利用来自卢旺达官方公报的约18000个释义对,设置三个温度尺度;第二阶段对715个经NLLB翻译的MNLI三元组进行微调,以学习蕴含结构;第三阶段使用英语-卢旺达语的OPUS-100翻译对对齐表示;第四阶段用从KinyaCOMET中筛选出的2936个高质量对进行优化,质量阈值为0.8。我们在SemRel2024-rw上进行评估,并引入了Wiki-RW-STS,这是一个新的无污染卢旺达语语义文本相似度(STS)基准,包含300个来自卢旺达语维基百科的对。一个七检查点集成模型(all5+23A*2,最终阶段权重加倍)在SemRel2024-rw上达到Spearman ρ=0.7298,比mE5-large高出20.9%,比OpenAI text-embedding-3-large高出41.0%。KinyaEmbed还在所有评估模型中取得了最佳的文档聚类轮廓系数(0.2146)。所有检查点、KinyaCOMET筛选的对以及Wiki-RW-STS基准均已公开提供。

英文摘要

We present KinyaEmbed, the first dedicated sentence embedding model for Kinyarwanda, a morphologically rich Bantu language spoken by over 12 million people in Rwanda. Existing multilingual embedding models such as LaBSE, mE5-large, and OpenAI text-embedding-3-large perform poorly on Kinyarwanda due to severe under-representation in their pre-training corpora. KinyaEmbed is built on KinyaBERT-large and trained via a four-stage curriculum using MultipleNegativesRankingLoss (MNRL): Stage 1 leverages ~18,000 paraphrase pairs from the Official Gazette of Rwanda with three temperature scales; Stage 2 fine-tunes on 715 NLLB-translated MNLI triplets for entailment structure; Stage 3 aligns representations using English-Kinyarwanda OPUS-100 translation pairs; Stage 4 refines with 2,936 high-quality pairs filtered from KinyaCOMET at quality threshold 0.8. We evaluate on SemRel2024-rw and introduce Wiki-RW-STS, a new contamination-free Kinyarwanda STS benchmark of 300 pairs derived from Kinyarwanda Wikipedia. A seven-checkpoint ensemble (all5+23A*2, with the final stage double-weighted) achieves Spearman \r{ho}=0.7298 on SemRel2024-rw, surpassing mE5-large by 20.9% and OpenAI text-embedding-3-large by 41.0%. KinyaEmbed also achieves the best document clustering silhouette score (0.2146) across all evaluated models. All checkpoints, the KinyaCOMET filtered pairs, and the Wiki-RW-STS benchmark are publicly available.

Comments14 pages, 3 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑