新加坡英语,行不行?针对新加坡英语的零样本语音合成的微调与评估
Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English
浏览论文内容
中文总结 AI 辅助
研究针对新加坡英语的零样本语音合成,对Chatterbox和CosyVoice 3两个模型在IMDA语料库上微调,评估三种语音分布在自然度等四维度的表现,发现微调能提升口音相似度,使生成分布更接近真实新加坡英语,是该领域首次系统研究。
中文摘要 AI 辅助
零样本语音到语音转换(ZS-TTS)在标准英语方面达到了接近人类的质量,但在模仿地区口音方面表现不佳。在一段简短的新加坡英语话语的提示下,最先进的系统在重现说话者音色的同时,将口音向通用英语扁平化。我们研究了对现成的ZS-TTS进行有针对性的微调是否能缩小新加坡英语(Singlish)的差距。我们在IMDA国家语音语料库中的50名新加坡英语使用者上对两个前沿的ZS-TTS模型Chatterbox和CosyVoice 3进行了微调。评估了三种语音分布:真实录音与由相同新加坡英语音频提示驱动的现成和微调生成的语音。评估涵盖四个维度:自然度、可懂度、说话者相似度和口音相似度。我们将适应(微调期间看到的域内说话者)与一致性(保留的说话者)分开,以测试口音转移是否能推广到训练数据之外。微调提高了Chatterbox和CosyVoice 3在域内和域外说话者上的口音相似度。它使生成的分布明显向真实的新加坡英语移动,并且在保留的说话者上这种提升仍然存在。据我们所知,这是第一项关于带有新加坡英语口音的语音合成的系统研究。
英文摘要
Zero-shot text-to-speech (ZS-TTS) achieves near-human quality for standard English, but it copies regional accents poorly. Prompted with a short Singlish utterance, state-of-the-art systems reproduce a speaker's timbre while flattening the accent toward generic English. We investigate whether targeted fine-tuning off-the-shelf ZS-TTS can close the gap for Singapore English (Singlish). We fine-tune two cutting-edge ZS-TTS models, Chatterbox and CosyVoice 3, on 50 Singlish speakers from the IMDA National Speech Corpus. Three speech distributions are evaluated: real recordings against off-the-shelf and fine-tuned generation driven by the same Singlish audio prompts. The evaluation covers four dimensions: naturalness, intelligibility, speaker similarity, and accent similarity. We separate adaptation (in-domain speakers seen during fine-tuning) from consistency (held-out speakers) to test whether accent transfer generalises beyond the training data. Fine-tuning raises accent similarity on in-domain and out-of-domain speakers for both Chatterbox and CosyVoice 3. It moves the generated distribution measurably toward real Singlish, with the gain persisting on held-out speakers. To our knowledge, this is the first systematic study of Singlish-accented TTS.