发表机构
Alibaba ATH Token Foundry(阿里巴巴ATH Token铸造厂)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多轮交互中TTS风格自适应难、音色漂移的问题,提出交互式TTS框架,将风格决策建模为可执行指令,结合迭代拒绝采样微调和上下文感知直接偏好优化,在VStyle和SpeechParaling-Bench上超越现有模型。
AI 中文摘要
在多轮多模态交互中,动态说话风格自适应仍然是文本到语音(TTS)系统面临的主要挑战。现有的上下文感知TTS(CTTS)方法通常以端到端的方式将对话上下文映射为语音。这种隐式建模使得上下文风格决策难以监督,同时风格、音色和内容的纠缠常常导致指令跟随能力弱以及跨轮次的严重音色漂移。为克服这些局限,我们提出交互式TTS,一个动态、风格自适应的框架,用于生成上下文恰当且说话人一致的语音。交互式TTS通过将上下文风格决策显式建模为可执行指令来解耦该过程。为弥合风格决策与语音生成之间的差距,我们引入了迭代拒绝采样微调(Iterative RSFT)和上下文感知直接偏好优化(CADPO),这些方法显著增强了指令跟随能力,并使生成的语音与对话上下文对齐。大量实验表明,交互式TTS在VStyle和SpeechParaling-Bench上优于最先进的模型。演示可在该https URL获取。
英文摘要
Dynamic speaking style adaptation in multi-turn multimodal interaction remains a major challenge for text-to-speech (TTS) systems. Existing context-aware TTS (CTTS) methods typically map dialogue context to speech in an end-to-end manner. Such implicit modeling makes contextual style decisions difficult to supervise, while the entanglement of style, timbre, and content often leads to weak instruction-following and severe timbre drift across turns. To overcome these limitations, we propose Interactive TTS, a dynamic, style-adaptive framework for contextually appropriate and speaker-consistent speech generation. Interactive TTS decouples the process by explicitly modeling contextual style decisions as executable instructions. To bridge the gap between style decisions and speech generation, we introduce Iterative Rejection Sampling Fine-Tuning (Iterative RSFT) and Context-Aware Direct Preference Optimization (CADPO), which significantly enhance instruction-following and align the generated speech with conversational contexts. Extensive experiments demonstrate that Interactive TTS outperforms state-of-the-art models on VStyle and SpeechParaling-Bench. Demo is available at https://wjtian-wonderful.github.io/InteractiveTTS/