EditVoice:基于编辑流的变长非自回归零样本TTS与语音编辑
EditVoice: Variable-Length Non-Autoregressive Zero-Shot TTS and Speech Editing with Edit Flows
浏览论文内容
中文总结 AI 辅助
EditVoice是首个变长非自回归零样本TTS模型,利用编辑流实现语音内容与长度联合更新,统一TTS与编辑任务,在多个基准上性能优异。
中文摘要 AI 辅助
近期非自回归(NAR)零样本文本到语音(TTS)模型并行生成语音,但通常需要在生成前指定目标序列长度。我们提出EditVoice,据我们所知,这是首个变长NAR零样本TTS模型,它使用编辑流(Edit Flows)通过插入、删除和替换操作联合更新语音内容和序列长度。EditVoice采用语音填充训练,统一了零样本TTS和基于文本的语音编辑,并允许在推理时使用前缀和后缀语音提示放置。我们引入互补提示采样(CPS),以利用两种提示放置所引发的互补编辑流预测。我们进一步发现,EditVoice可以编辑超出其训练来源的源语音和模型生成语音。我们利用这种泛化能力进行端到端编辑和无需训练的后生成精炼。凭借在10K小时GigaSpeech上训练的编辑流模型,EditVoice在Seed-TTS Eval EN和LibriSpeech-PC上展示了具有竞争力的零样本TTS性能,并在RealEdit上展示了语音编辑性能。
英文摘要
Recent non-autoregressive (NAR) zero-shot text-to-speech (TTS) models generate in parallel but typically require the target sequence length to be specified before generation. We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Edit Flows to jointly update speech content and sequence length through insertions, deletions, and substitutions. EditVoice adopts speech-infilling training, which unifies zero-shot TTS and text-based speech editing and allows both prefix and suffix speech prompt placements at inference. We introduce Complementary Prompt Sampling (CPS) to leverage the complementary Edit Flow predictions induced by the two prompt placements. We further find that EditVoice can edit source and model-generated speech beyond its training sources. We use this generalization for end-to-end editing and training-free post-generation refinement. With the Edit Flow model trained on 10K h of GigaSpeech, EditVoice demonstrates competitive zero-shot TTS performance on Seed-TTS Eval EN and LibriSpeech-PC and speech editing performance on RealEdit.
发表机构
- School of Informatics, Xiamen University(厦门大学信息学院)
- School of Electronic Science and Engineering, Xiamen University(厦门大学电子科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。