arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29146cs.CL

BanglaKontho:弥合孟加拉语文本到语音转换中的长格式差距

BanglaKontho: Closing the Long-Form Gap in Bangla Text-to-Speech

Mizbaul Haque Maruf

首次发表
浏览论文内容

中文总结 AI 辅助

针对孟加拉语长格式TTS资源不足,提出20小时单说话人语料库BanglaKontho及文本规范化器,基线MB-iSTFT-VITS显著优于IndicTTS-Bn,并公开数据。

中文摘要 AI 辅助

孟加拉语是世界第七大使用语言,但在神经文本到语音转换方面资源仍然不足。公开的孟加拉语语音语料库主要由为语音识别收集的短朗读提示话语组成,未覆盖长格式韵律和一致的单说话人叙述。我们提出了BanglaKontho,一个由专业有声书录音衍生的20小时单说话人孟加拉语TTS语料库:包含7,050个分段话语,带有24 kHz下验证的转录文本。我们还发布了一个可复用的孟加拉语文本规范化器,涵盖孟加拉式数字分组、货币和日期表达、Danda标点和Unicode规范化,以及完整的预处理流程。从头训练的MB-iSTFT-VITS基线达到9.5%的词错误率和4.46的自然度MOS,而相同架构在12小时IndicTTS-Bn语料库上重新训练则分别为16.0%和3.16。该语料库在CC BY-NC 4.0许可下公开发布。

英文摘要

Bangla, the seventh most spoken language in the world, remains under-resourced for neural text-to-speech. Public Bangla speech corpora are dominated by short read-prompt utterances collected for speech recognition, leaving long-form prosody and consistent single-speaker narration uncovered. We present BanglaKontho, a single-speaker Bangla TTS corpus of 20 hours derived from professional audiobook recordings: 7,050 segmented utterances with verified transcripts at 24 kHz. We also release a reusable Bangla text normalizer covering Bangladeshi-style digit grouping, currency and date expressions, Danda punctuation and Unicode normalization, together with the full preprocessing pipeline. An MB-iSTFT-VITS baseline trained from scratch reaches 9.5% WER and 4.46 naturalness MOS, against 16.0% and 3.16 for the same architecture retrained on the 12-hour IndicTTS-Bn corpus. The corpus is released openly under CC BY-NC 4.0.

发表机构

  • Vivasoft Limited(Vivasoft有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑