arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08362eess.AScs.HCcs.SD

CtrlSpeech:面向富有表现力的语音合成的从粗到细控制

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

Zhisheng Zheng, Xiaohang Sun, Zhu Liu, Caren Chen, Rohith Kumar, Manoj Aggarwal, Gerard Medioni, David Harwath

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有TTS系统难以实现词/音素级细粒度表达控制的问题,提出基于DiTAR架构的CtrlSpeech框架,结合全局说话人条件与音素对齐韵律信号,实现灵活可控的表达语音合成,兼具零样本性能与更优可控性。

中文摘要 AI 辅助

近期的文本转语音(TTS)系统已实现出色的自然度与零样本语音克隆性能,但在词或音素层面对富有表现力的语音进行细粒度控制仍具挑战性。我们提出CtrlSpeech,这是一种具备从粗到细控制能力的可控且富有表现力的TTS框架,构建于DiTAR架构之上,结合全局说话人条件与音素对齐的基频、响度及时长信号,在保留目标说话人音色的同时实现局部韵律控制。该设计使用户能以精细时间粒度调整表达属性,使语音优化更灵活可控。实验结果表明,CtrlSpeech实现了有竞争力的零样本TTS性能,并提升了表达属性的可控性,证明其在灵活实用的富有表现力的语音合成方面的有效性。

英文摘要

Recent Text-To-Speech (TTS) systems have achieved strong naturalness and zero-shot voice cloning performance, but fine-grained control of expressive speech at the word or phoneme level remains challenging. We propose CtrlSpeech, a controllable, expressive TTS framework with coarse-to-fine control. Built on the DiTAR architecture, CtrlSpeech combines global speaker conditioning with phone-aligned pitch, loudness, and duration signals, enabling localized prosodic control while preserving the target speaker's timbre. This design allows users to adjust expressive attributes at a fine temporal granularity, making speech refinement more flexible and controllable. Experimental results show that CtrlSpeech achieves competitive zero-shot TTS performance and improves controllability over expressive attributes, demonstrating its effectiveness for flexible and practical expressive speech synthesis.

发表机构

  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • Amazon(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑