发表机构
University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对班图声调语言文本到语音合成中声调信息缺失的问题,提出Morpho-VITS模型,通过语素序列编码器和音素到语素注意力网络显式建模形态句法,在基尼亚卢旺达语上显著提升合成语音的自然度、语调和可懂度。
AI 中文摘要
班图声调语言的文本到语音模型面临着一个源于词汇(即词、词干和词缀的清单)和语法(即形态句法)的声调系统的挑战。使问题更加复杂的是,这些语言的标准书写系统常常省略声调标记和音节时长信息,读者必须根据上下文来消除歧义。受班图语言声调系统的语言学描述的启发,我们提出了一种端到端的文本到语音模型,该模型在文本编码机制中增加了形态句法先验。我们将VITS架构中的标准音素编码器替换为语素序列编码器和音素到语素的注意力网络。我们假设,通过使用这种显式的形态学建模,我们可以捕获生成正确声调所需的信息。在基尼亚卢旺达语(一种声调且形态复杂的班图语言)上进行的实验表明,这种形态学建模显著改善了文本到语音合成。具体来说,所提出的方法显著提高了所生成合成语音的自然度、语调性和可懂度。
英文摘要
Text-to-speech models for Bantu tonal languages are challenged by a tonal system that is rooted in both the lexis (i.e., the inventory of words, stems, and affixes) and the grammar (i.e., morpho-syntax). To complicate matters, the standard writing systems of these languages often omit tone markings and syllable duration information, which must be disambiguated by the reader based on context. Motivated by linguistic descriptions of Bantu language tone systems, we propose an end-to-end text-to-speech model that augments the text encoding mechanism with a morpho-syntactic prior. We replace the standard phoneme encoder in the VITS architecture with a morpheme sequence encoder and a phoneme-to-morpheme attention network. We posit that, by using this explicit morphological modeling, we can capture the information required to produce the correct tone. Experiments conducted on the Kinyarwanda language, a tonal and morphologically complex Bantu language, reveal substantial TTS improvement from this morphological modeling. Specifically, the proposed method significantly improves the naturalness, intonation, and intelligibility of the produced synthetic voices.
Comments5 pages, 2 figures, 2 tables