CtrlSpeech:面向富有表现力的语音合成的从粗到细控制
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis
浏览论文内容
中文总结 AI 辅助
针对现有TTS系统难以实现词/音素级细粒度表达控制的问题,提出基于DiTAR架构的CtrlSpeech框架,结合全局说话人条件与音素对齐韵律信号,实现灵活可控的表达语音合成,兼具零样本性能与更优可控性。
中文摘要 AI 辅助
近期的文本转语音(TTS)系统已实现出色的自然度与零样本语音克隆性能,但在词或音素层面对富有表现力的语音进行细粒度控制仍具挑战性。我们提出CtrlSpeech,这是一种具备从粗到细控制能力的可控且富有表现力的TTS框架,构建于DiTAR架构之上,结合全局说话人条件与音素对齐的基频、响度及时长信号,在保留目标说话人音色的同时实现局部韵律控制。该设计使用户能以精细时间粒度调整表达属性,使语音优化更灵活可控。实验结果表明,CtrlSpeech实现了有竞争力的零样本TTS性能,并提升了表达属性的可控性,证明其在灵活实用的富有表现力的语音合成方面的有效性。
英文摘要
Recent Text-To-Speech (TTS) systems have achieved strong naturalness and zero-shot voice cloning performance, but fine-grained control of expressive speech at the word or phoneme level remains challenging. We propose CtrlSpeech, a controllable, expressive TTS framework with coarse-to-fine control. Built on the DiTAR architecture, CtrlSpeech combines global speaker conditioning with phone-aligned pitch, loudness, and duration signals, enabling localized prosodic control while preserving the target speaker's timbre. This design allows users to adjust expressive attributes at a fine temporal granularity, making speech refinement more flexible and controllable. Experimental results show that CtrlSpeech achieves competitive zero-shot TTS performance and improves controllability over expressive attributes, demonstrating its effectiveness for flexible and practical expressive speech synthesis.
发表机构
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
- Amazon(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。