发表机构
University of Edinburgh; Huawei; Wallenberg AI, Autonomous Systems and Software Program(爱丁堡大学; 华为; 瓦伦堡人工智能、自主系统与软件计划)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对低资源下高表现力TTS的标签稀缺问题,提出迭代自学习框架,基于Invert-Classify方法迭代优化伪标签,在词级突出性和情感任务上验证,提升了标签贴合度与合成质量,性能接近全监督模型。
AI 中文摘要
采用显式条件标签的高表现力文本到语音(TTS)系统可对表现力属性提供直接且可解释的控制,这与基于参考或提示的方法形成对比,但此类系统需要带标签的数据。大规模获取这些标签成本高昂且耗时,然而此前没有半监督框架针对这一特定瓶颈问题。现有的半监督TTS方法反而针对的是成对语音-文本数据或转录本的稀缺性。为解决表现力标签的稀缺问题,我们提出了一种用于高表现力TTS的迭代自学习(ISL)框架,该框架基于Invert-Classify构建,这是一种无分类器的方法,可通过反转冻结的生成模型来恢复离散的表现力标签。该框架迭代地使用当前模型对未标记语音进行伪标签标注,在标记数据与伪标记数据的组合上重新训练,然后重复该过程,逐步优化标签质量和合成效果。我们在两项表现力任务上进行了验证,即词级突出性和话语级情感,覆盖了多个低资源数据分割。我们发现,迭代优化可提升伪标签准确率,优于单次基线方法。此外,我们观察到,这些表现力伪标签的改进转化为表现力标签贴合度和合成质量的提升,通过客观指标和人工听觉测试得到了验证。在数据最稀缺的条件下,经ISL训练的模型优于单次伪标签方法,且进一步接近全监督性能,表明基于梯度的ISL是低资源TTS中解决表现力标签稀缺问题的有效方案。
英文摘要
Expressive text-to-speech (TTS) systems that use explicit conditioning labels provide direct and interpretable control over expressive attributes, in contrast to reference-based or prompting-based approaches, but require labeled data. Obtaining these labels at scale is costly and time-consuming, yet no prior semi-supervised framework addresses this specific bottleneck. Existing semi-supervised TTS methods instead target scarcity of paired speech-text data or transcriptions. To address the scarcity of expressive labels, we propose an Iterative Self-Learning (ISL) framework for expressive TTS, built on Invert-Classify, a classifier-free method that recovers discrete expressive labels by inverting a frozen generative model. The framework iteratively pseudo-labels unlabeled speech using the current model, retrains on the combined labeled and pseudo-labeled data, and repeats, progressively refining label quality and synthesis. We validate on two expressive tasks, word-level prominence and utterance-level emotion, across multiple low-resource data splits. We find that iterative refinement can improve pseudo-label accuracy over single-pass baselines. Furthermore, we observe that these improvements in pseudo-labeling of expressivity translate to gains in expressive label adherence and synthesis quality, confirmed by objective metrics and human listening tests. In the most data-scarce conditions, ISL-trained models outperform single-pass pseudo-labeling and further approach fully supervised performance, demonstrating that gradient-based ISL is an effective solution to expressive label scarcity in low-resource TTS.