发表机构
Institute for Language and Speech Processing, Athena R.C.; Department of Digital Medicine, University of Bern; School of Electrical and Computer Engineering, National Technical University of Athens(雅典娜研究与创新中心语言与语音处理研究所; 伯尔尼大学数字医学系; 雅典国家技术大学电气与计算机工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出数据整理与确定性提示及LoRA微调方案,在3.5小时数据上实现低资源希腊语TTS,达到WER 10.7%和接近人类的说话人一致性。
AI 中文摘要
现代TTS系统在高资源语言上接近人类质量,但在干净语音数据稀缺时会退化。现代希腊语就是这种情况,缺乏最先进合成背后的精选语料库。我们提出了一种数据整理方案,通过WhisperX对齐和过滤将有声读物录音转换为TTS就绪数据。然后我们微调Parler-TTS(880M),这是一个基于提示的多语言模型,其预训练编码了可迁移到希腊语的语音先验。在开发过程中,我们发现LLM生成的风格提示在推理时引入说话人漂移。用确定性提示替换它们解决了这一问题,并且在3.5小时单说话人数据上训练说话人特定的LoRA阶段锚定了身份,同时仅更新约5%的参数。我们的系统实现了WER 10.7%(高于ASR下限2.9),MOS-I 4.00(对比人类语音4.36),以及接近人类的说话人一致性(MOS-C 4.24对比4.30),表明使用有限的精选数据可以实现稳健的单说话人希腊语TTS。
英文摘要
Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, lacking the curated corpora behind state-of-the-art synthesis. We propose a data curation recipe that transforms audiobook recordings into TTS-ready data via WhisperX alignment and filtering. Then we fine-tune Parler-TTS (880M), a prompt-based multilingual model whose pre-training encodes phonetic priors transferable to Greek. During development, we find that LLM-generated style prompts introduce speaker drift at inference. Replacing them with deterministic prompts resolves this, and a speaker-specific LoRA stage trained on 3.5 h of single-speaker data anchors identity while updating ~5% of parameters. Our system achieves WER 10.7% (2.9 above the ASR floor), MOS-I 4.00 (vs. 4.36 human speech), and near-human speaker consistency (MOS-C 4.24 vs. 4.30), showing that robust single-speaker Greek TTS is achievable with limited curated data.
CommentsInterspeech 2026