DEFINE:基于范例引导的口音控制用于零样本语音合成
DEFINE: Exemplar-Guided Accent Control for Zero-Shot TTS
浏览论文内容
中文总结 AI 辅助
DEFINE通过解耦说话人身份与口音,利用范例编码器和引导权重,在零样本TTS中实现口音独立控制,提升口音准确率并泛化至未见口音。
中文摘要 AI 辅助
零样本文本到语音(TTS)能够从短参考录音中重现未见过的说话人,但通常会将说话人身份和口音纠缠在同一参考中。我们提出DEFINE,一个端到端框架,通过将说话人身份和目标口音分别条件于不同的音频范例来解耦这些因素。单一的推理时引导权重可连续控制口音强度,无需重新训练。DEFINE基于F5-TTS构建,采用参数高效的LoRA适配,通过范例编码器将短口音范例映射到条件空间,该编码器通过学习的口音原型进行监督,推理时既不需要口音标签,也不需要合成后的波形转换。在已见口音上,增加口音引导将口音探测准确率从6.5%提升至19.6%。更重要的是,单个DEFINE模型能够将其口音控制泛化到训练口音集之外:在已见和域外口音上(尽管不在保留口音上),它匹配了双模型TTS-语音转换级联的口音迁移性能,同时实现了更高的说话人相似度和相当的可预测语音质量。这些结果表明,在单个零样本TTS模型中,说话人身份和口音可以从音频范例中独立控制,包括对训练期间未见过的口音。
英文摘要
Zero-shot text-to-speech (TTS) can reproduce an unseen speaker from a short reference recording, but typically entangles speaker identity and accent within the same reference. We introduce DEFINE, an end-to-end framework that decouples these factors by conditioning speaker identity and target accent on separate audio exemplars. A single inference-time guidance weight continuously controls accent strength without retraining. Built on F5-TTS with parameter-efficient LoRA adaptation, DEFINE maps short accent exemplars into a conditioning space using an exemplar encoder supervised through learned accent prototypes, requiring neither accent labels at inference time nor post-synthesis waveform conversion. On seen accents, increasing accent guidance improves accent-probe accuracy from 6.5% to 19.6%. More importantly, a single DEFINE model generalizes accent control beyond its training accent set: on seen and out-of-domain accents, though not on held-out accents, it matches the accent transfer performance of a two-model TTS-voice-conversion cascade while achieving higher speaker similarity and comparable predicted speech quality. These results demonstrate that speaker identity and accent can be independently controlled from audio exemplars within a single zero-shot TTS model, including for accents unseen during training.