arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越语音字幕:面向对话式文本到语音的语音奖励风格规划

Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

Shiao Zhu, Lianbo Liu, Sizhen Lyu, Yuzhe Wang, Sheng Li, Takahiro Shinozaki

arXiv 2610.11461首次发表:更新:

发表机构

Institute of Science Tokyo(东京科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对对话式TTS提出SRSP方法,以冻结TTS模型为基础,用GRPO优化风格规划器,在ISCSLP 2026 CoT-TTS语料库子集上,其语音风格、情感相似度及语境表现均优于基线。

AI 中文摘要

自然语言风格描述为大语言模型(LLMs)与可控文本到语音(TTS)提供了可解释的接口。然而,将描述作为伪标签会将目标声学特征压缩为文本,描述的保真度并不意味着能有效控制特定合成器。我们通过实验表明,对于同一话语的候选指令,语音-文本对齐仅能微弱预测下游声学相似度。因此,我们提出语音奖励风格规划(SRSP),该方法通过冻结的下游TTS模型训练基于文本的风格规划器。给定对话历史和响应文本,规划器生成候选指令,并使用组相对策略优化(GRPO)进行优化,以目标语音标记的教师强制似然作为奖励。在ISCSLP 2026 CoT-TTS语料库的英语子集上,SRSP相比基础LLM和目标音频感知字幕基线,实现了更高的语音风格和情感相似度,以及更低的梅尔倒谱失真。基于LLM的表达性语音评估进一步显示,SRSP在语境恰当性和参考一致性上优于所有基线。

英文摘要

Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer. We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance. We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model. Given dialogue history and response text, the planner generates candidate instructions and is optimized with group-relative policy optimization (GRPO), using the teacher-forced likelihood of target speech tokens as the reward. On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines. LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑