发表机构
Institute of Science Tokyo(东京科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对对话式TTS提出SRSP方法,以冻结TTS模型为基础,用GRPO优化风格规划器,在ISCSLP 2026 CoT-TTS语料库子集上,其语音风格、情感相似度及语境表现均优于基线。
AI 中文摘要
自然语言风格描述为大语言模型(LLMs)与可控文本到语音(TTS)提供了可解释的接口。然而,将描述作为伪标签会将目标声学特征压缩为文本,描述的保真度并不意味着能有效控制特定合成器。我们通过实验表明,对于同一话语的候选指令,语音-文本对齐仅能微弱预测下游声学相似度。因此,我们提出语音奖励风格规划(SRSP),该方法通过冻结的下游TTS模型训练基于文本的风格规划器。给定对话历史和响应文本,规划器生成候选指令,并使用组相对策略优化(GRPO)进行优化,以目标语音标记的教师强制似然作为奖励。在ISCSLP 2026 CoT-TTS语料库的英语子集上,SRSP相比基础LLM和目标音频感知字幕基线,实现了更高的语音风格和情感相似度,以及更低的梅尔倒谱失真。基于LLM的表达性语音评估进一步显示,SRSP在语境恰当性和参考一致性上优于所有基线。
英文摘要
Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer. We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance. We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model. Given dialogue history and response text, the planner generates candidate instructions and is optimized with group-relative policy optimization (GRPO), using the teacher-forced likelihood of target speech tokens as the reward. On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines. LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.