弥合数据、推理与对齐:面向上下文感知指令跟随TTS的统一框架
Bridging Data, Reasoning, and Alignment: A Unified Framework for Context-Aware Instruction-Following TTS
浏览论文内容
中文总结 AI 辅助
针对CoT-TTS挑战,提出包含数据清洗、CA-DPO对齐及LLM评估的统一优化框架,显著提升上下文感知指令跟随语音合成的推理与执行一致性。
中文摘要 AI 辅助
ISCSLP 2026 CoT-TTS挑战赛要求TTS系统在合成上下文适当的语音之前,先从对话历史中生成思维链(CoT)推理。虽然官方基线建立了统一的架构,但其仍受限于上下文理解能力有限、指令保真度弱以及音频质量欠佳等问题。我们提出了一套系统性的优化流程来解决这些局限性。首先,我们开发了一个数据处理框架,通过FullSubNet去噪、Qwen3-ASR重新转录以及基于Qwen3.5-35B-A3B的历史-CoT一致性分析来清洗原始数据,同时利用Qwen3-TTS和Seed-VC在严格质量过滤下蒸馏出54.5万条高保真指令样本。其次,我们提出了一种上下文感知直接偏好优化(CA-DPO)方法。通过采用级联过滤策略、ASR预筛选、LLM锦标赛排序和说话人相似性验证,我们获得了高置信度的偏好对,显著增强了DPO训练期间整体的“上下文→CoT→语音”一致性。第三,我们建立了一种评估方法,包含一个500样本的测试集和一个LLM-as-Judge框架,以独立评估推理和执行保真度。实验表明,我们的系统在所有客观和主观指标上均显著优于基线,验证了我们的数据治理和对齐策略。
英文摘要
The ISCSLP 2026 CoT-TTS Challenge requires TTS systems to generate Chain-of-Thought (CoT) reasoning from dialogue history before synthesizing contextually appropriate speech. While the official baseline establishes a unified architecture, it remains constrained by limited contextual comprehension, weak instruction fidelity, and suboptimal audio quality. We present a systematic optimization pipeline to address these limitations. First, we develop a data process framework that cleans raw data via FullSubNet denoising, Qwen3-ASR re-transcription, and Qwen3.5-35B-A3B-based history-CoT consistency analysis, while distilling 545K high-fidelity instruction samples using Qwen3-TTS and Seed-VC under strict quality filtration. Second, we propose a Context-Aware Direct Preference Optimization (CA-DPO) method. By employing a cascaded filtering strategy, ASR prescreening, LLM tournament ranking, and speaker similarity verification, we obtain high-confidence preference pairs that significantly enhance holistic ``Context$\rightarrow$CoT$\rightarrow$Speech'' consistency during DPO training. Third, we establish an evaluation method featuring a 500-sample test set and an LLM-as-Judge framework to independently assess reasoning and execution fidelity. Experiments demonstrate that our system significantly outperforms the baseline across all objective and subjective metrics, validating our data governance and alignment strategies.
发表机构
- Northwestern Polytechnical University(西北工业大学)
- Shenzhen Pimei Technology Co., Ltd.(深圳市湃美科技有限公司)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。