arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

COT-TTS:基于思维链推理的音频上下文感知文本到语音合成

COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning

Weizhen Bian, Sitong Cheng, Rongxiu Zhong, Jiahao Pan, Liumeng Xue, Boyi Kang, Shilei Zhang, Jinglei Liu, Yue Wang, Junlan Feng, Bei Liu, Wei Xue

arXiv 2609.22697首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; JIUTIAN Research, China Mobile; The State Key Laboratory of Multimedia Information Processing, Peking University; Nanjing University; China Mobile (Hong Kong) Innovation Research Institute(香港科技大学; 中国移动九天研究院; 北京大学多媒体信息处理国家重点实验室; 南京大学; 中国移动(香港)创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出COT-TTS任务,基于对话上下文推理实现文本到语音合成,构建大规模双语数据集和基线,开发0.6B/1.7B自回归模型,以较少参数达到与大规模基线相当的性能,并支持情感、重音和节奏变化。

AI 中文摘要

近年来,文本到语音合成系统在语音表现力和可控性方面取得了显著进展。然而,生成语音的说话风格通常依赖于用户明确指定的指令。在自然对话中,说话风格应从先前的对话上下文中自然推断出来。因此,我们提出了COT-TTS,一个基于上下文感知和推理的文本到语音合成任务。给定历史对话音频、目标文本和参考语音,系统应理解对话上下文,推断出明确的中间推理,并最终以指定的音色合成目标语音。为支持该任务,我们构建了一个大规模的双语对话语音数据集,包含900万个训练样本,其中包括100万个样本的高质量子集。我们还构建了一个源分离的基准测试集,包含800个人工验证的样本,并建立了强大的任务特定基线。此外,我们开发了参数规模为0.6B和1.7B的端到端自回归模型,生成情感标注的转录文本、可编辑的语音风格推断和语音标记。实验结果表明,所提出的模型在参数显著减少的情况下,达到了与大规模基线系统相当的性能。同时,该模型在时长一致性和情感一致性方面表现良好,并能根据对话上下文生成适当的情感、重音和节奏变化。为促进未来研究,我们将公开发布数据构建流程、数据集、训练模型及相关资源。演示页面和其他资源可在以下https URL获取。

英文摘要

Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated speech typically relies on clear user-specified instructions. In natural conversations, speaking style should be naturally inferred from the preceding conversational context. Therefore, we propose COT-TTS, a context-aware, reasoning-based text-to-speech task. Given historical conversation audio, target text, and a reference speech, the system should comprehend the conversational context, infer an explicit intermediate reasoning, and finally synthesize the target speech with the specified timbre. To support this task, we constructed a large-scale bilingual conversational speech dataset comprising 9 million training samples, including a high-quality subset of 1 million samples. We further constructed a source-disjoint benchmark with 800 human-verified samples and established strong task-specific baselines. Additionally, we developed end-to-end autoregressive models with parameter sizes of 0.6B and 1.7B, generating emotion-labeled transcripts, editable speech style inferences, and speech tokens. Experimental results show that the proposed model achieves performance comparable to large-scale baseline systems with significantly fewer parameters. At the same time, the model performs well in terms of duration consistency and emotional consistency, and can generate appropriate emotional, stress, and rhythmic variations based on the conversational context. To facilitate future research, we will publicly release the data construction pipeline, dataset, trained models, and related resources. The demo page and additional resources are available at https://luckybian.github.io/COT-TTS

CommentsUnder review at IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑