AI 中文总结
本文提出BanglaBox,一个语音和性别平衡的孟加拉语语料库及三项微调改进,实现英语预训练TTS模型对孟加拉语的高效适配,在零样本声音克隆中达到接近自然的自然度并优于基线,且微调音频减少约7倍。
AI 中文摘要
我们提出了一种将英语预训练的自回归文本到语音(TTS)基础模型适配到代表性不足语言的方案,并以孟加拉国孟加拉语为例进行了演示。现有的孟加拉语TTS语料库规模小且为单说话人,据我们所知,目前尚无针对孟加拉国口音的开源零样本声音克隆系统。我们贡献了一个语音和性别平衡的两层孟加拉国孟加拉语语料库,通过基于连接簇(juktakkhor)的分层Jensen-Shannon散度目标进行平衡,同时提出了三项微调改进:合并一致的词元化器扩展、孟加拉语文本规范化以及提示掩蔽的双损失目标。这些改进在语言切换过程中保留了零样本克隆能力。我们的BanglaEval协议对母语者的评分应用了带有Bonferroni校正的Wilcoxon符号秩检验。BanglaBox在自然度上接近自然水平,在自然度、说话人相似度和清晰度上优于商业和开源基线,并且在使用大约7倍少的孟加拉语微调音频的情况下,达到了与先前少样本结果相当的说话人相似度。我们进一步使用包含七类困难真实世界文本的压力测试集、公开的BnTTS评估基准以及自然出现的(非语言模型生成的)孟加拉语,验证了该方案在我们自己的测试集之外的有效性。所有工件,包括语料库、权重、代码和完整的评估材料,均已公开并无条件发布。
英文摘要
We present a recipe for adapting English-pretrained autoregressive TTS foundation models to underrepresented languages, demonstrated on Bangladeshi Bangla. Existing Bangla TTS corpora are small and single-speaker, and to our knowledge no open zero-shot voice-cloning system is available for the Bangladeshi register. We contribute a phonetically- and gender-balanced two-tier Bangladeshi Bangla corpus balanced via a tiered Jensen-Shannon divergence objective over conjunct clusters (juktakkhor), together with three fine-tuning changes: a merge-consistent tokenizer extension, Bangla text normalization, and a prompt-masked dual-loss objective. These changes preserve zero-shot cloning across the language switch. Our BanglaEval protocol applies Wilcoxon signed-rank tests with Bonferroni correction over native-speaker ratings. BanglaBox attains near-natural Naturalness, outperforms commercial and open-source baselines on Naturalness, Speaker Similarity, and Clarity, and reaches speaker similarity comparable to prior few-shot results while using approximately 7x less Bangla fine-tuning audio. We further validate the recipe beyond our own test split using a seven-category stress set of difficult real-world text, the public BnTTS evaluation benchmarks, and naturally occurring Bangla that no language model wrote. All artifacts, including the corpus, weights, code, and complete evaluation materials, are released publicly and unconditionally.
Comments24 pages, 6 figures, 25 tables, and 1 algorithm. Includes appendices and the ACL Responsible NLP Checklist. Corpus, model weights, code, and evaluation materials: https://huggingface.co/Banglabox