ÌròyìnSpeech 文本语料库:用于语音和语言技术的 24,905 条精选约鲁巴语句子
The ÌròyìnSpeech Text Corpus: 24,905 Curated Yorùbá Sentences for Speech and Language Technology
浏览论文内容
中文总结 AI 辅助
本文发布 ÌròyìnSpeech 文本语料库,包含 24,905 条人工验证的约鲁巴语带声调句子,支持语音技术研究,并记录了 Unicode 规范化问题及修正。
中文摘要 AI 辅助
ÌròyìnSpeech 是一个 42 小时、80 位说话人的约鲁巴语朗读语音语料库,其音频自 2024 年起由 ELRA 分发。本文描述了其文本组件的发布:24,905 条独特的、人工验证的、带声调标记的约鲁巴语句子(275,897 个词元;15,687 个词型),这些句子于 2022 年作为录音提示语整理而成。其中约 11,000 条句子改编自开放许可的新闻材料;其余句子为内部撰写,以扩大覆盖面,超越现有约鲁巴语语料库中占主导地位的宗教翻译内容。每个句子都经过人工检查声调标记的准确性,并编辑以确保朗读清晰、语域中性,且经过本地化处理,使非约鲁巴语的人名和地名以约鲁巴语形式出现。准备发布文本时,发现了影响超过 60% 行(同一字母的预组合形式和分解形式在单个句子中同时出现)的系统性 Unicode 规范化失败,我们对此进行了记录和纠正。该语料库支持变音符号恢复、字素到音素转换、TTS 前端开发和正字法研究,并可作为新录音的验证提示集。
英文摘要
ÌròyìnSpeech is a 42-hour, 80-speaker Yorùbá read-speech corpus whose audio has been distributed by ELRA since 2024. This paper describes the release of its text component: 24,905 unique, hand-verified, tone-marked Yorùbá sentences (275,897 tokens; 15,687 types), curated in 2022 as recording prompts. Roughly 11,000 sentences were adapted from openly licensed news material; the remainder were written in-house to broaden coverage beyond the religious translation that dominates existing Yorùbá corpora. Every sentence was checked by hand for tone-mark accuracy, edited for read-aloud clarity and a neutral register, and localised so that non-Yorùbá personal and place names appear in Yorùbá form. Preparing the text for release surfaced systematic Unicode normalisation failures affecting more than 60% of lines (with precomposed and decomposed forms of the same letter co-occurring within single sentences) which we document and correct. The corpus supports diacritic restoration, grapheme-to-phoneme conversion, TTS front-end development and orthographic research, and serves as a validated prompt set for new recording.
发表机构
- Stanford University(斯坦福大学)
- Niger-Volta LTI
- Mila / McGill University(Mila / 麦吉尔大学)
机构由 AI 辅助整理,请以论文原文为准。