arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TranslatePsy-AfriSLM:面向低资源机器翻译的高质量数据扩展

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

Milan Gritta, Patrik Lambert, Jihye Back, Amril Nazir

arXiv 2608.18655首次发表:更新:

发表机构

Tether AI Research(泰瑟人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对非洲语言机器翻译的数字鸿沟问题,推出面向19种撒哈拉以南非洲语言的开源资源TranslatePsy-AfriSLM,经质量过滤后构建的小参数模型性能远超更大规模的同类系统。

AI 中文摘要

人工智能的快速发展在很大程度上忽略了非洲语言,造成了数字鸿沟,限制了非洲大陆对人工智能的应用。近期的开源大语言模型(LLM)在非洲语言机器翻译任务上系统性表现不佳,而大规模、高质量的开源平行数据的缺乏,制约了具有竞争力的小语言模型(SLM)的发展。我们推出了TranslatePsy-AfriSLM,这是一套面向19种撒哈拉以南非洲语言的开源机器翻译资源,包括经过筛选的平行数据、非洲专属合成数据,以及一系列微调后的小语言模型。我们的实证研究表明,统一质量评估过滤可去除多达96%的训练 token 而不降低质量,且过滤后的合成数据在质量-效率帕累托前沿中占据主导地位。基于所得数据混合物进行微调后,仅含0.8B参数的TranslatePsy-AfriSLM在性能上远超规模大得多的系统,包括TranslateGemma-27B和Qwen3.5-122B-A10B。

英文摘要

The rapid progress in Artificial Intelligence has largely bypassed African languages, creating a digital divide that limits AI adoption on the continent. Recent open-source LLMs systematically underperform on African language machine translation, while the lack of large-scale, high-quality, open-source parallel data has constrained the development of competitive small language models (SLMs). We introduce TranslatePsy-AfriSLM, a collection of open-source machine translation resources for 19 Sub-Saharan African languages, including curated parallel data, African-specialized synthetic data, and a family of fine-tuned SLMs. Our empirical study shows that unified quality-estimation filtering removes up to 96% of training tokens without degrading quality, and that filtered synthetic data dominates the quality-efficiency Pareto frontier. Fine-tuned on the resulting data mixture, TranslatePsy-AfriSLMs outperform substantially larger systems, including TranslateGemma-27B and Qwen3.5-122B-A10B, with as few as 0.8B parameters.

CommentsEMNLP 2026 (main conference)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑