Yo-ByT5:约鲁巴语的高效高保真变音符号恢复
Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá
- University of Lagos(拉各斯大学)
- Bayero University Kano(巴耶罗大学卡诺分校)
- Imperial College London(伦敦帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出Yo-ByT5,一种从ByT5-small微调的字节级自动变音符号恢复模型,在YAD基准上以更少参数达到与mT5-base相当的性能,并展现更高文本保真度,同时发布代码并呼吁构建更大基准。
AI中文摘要:
约鲁巴语是一种广泛使用的声调语言,依赖变音符号来避免词汇歧义。然而,该语言经常在没有这些变音符号的情况下书写,从而阻碍了下游自然语言处理(NLP)任务。在本文中,我们介绍了Yo-ByT5,一种从ByT5-small微调而来的字节级自动变音符号恢复(ADR)模型。我们在一致协议下,将Yo-ByT5与五个公开发布的约鲁巴语ADR模型和一个开放权重的大型语言模型(LLM)在YAD基准上进行了评估。我们的结果表明,Yo-ByT5在DER为10.14%和CER为3.48%的情况下,与现有最强模型mT5-base的性能相当。此外,尽管其参数量约为mT5-base的一半,它仍展现出优越的文本保真度。我们还发布了训练代码和模型输出,并呼吁开发一个更大、专门为约鲁巴语变音符号恢复设计的基准。
英文摘要:
Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.