arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28970cs.SDcs.AI

诊断后优化:基于AudioLLM引导校正的闭环TTS系统

Diagnose, Then Refine: A Closed-Loop TTS System with AudioLLM-Guided Correction

  • National University of Singapore(新加坡国立大学)
  • Tencent(腾讯)
  • LIGHTSPEED(光速)
  • Nanyang Technological University(南洋理工大学)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Shenzhen Loop Area Institute(深圳河套学院)

机构由 AI 辅助整理,请以论文原文为准。

Zeyang Song, Tianchi Liu, Tianrui Wang, Chenglin Xu, Steven Y. Guo, Haizhou Li

AI总结:

针对开环TTS易出现韵律缺陷且指标难检测的问题,提出含AudioLLM评判器与细粒度优化器的LoopTTS框架,构建4.2万示例数据集训练优化器,实验表明其修复质量与指令遵循能力更优。

AI中文摘要:

当前TTS系统通常依赖开环单次生成,易产生零星局部韵律缺陷,如重音错位、停顿不自然或语调平淡,而 utterance-level 指标往往无法检测到这些问题。我们提出LoopTTS,一种由评判器引导的Filter-Judge-Refiner框架,用于修复经AudioLLM诊断的低质量TTS输出。从基础TTS模型生成初始话语后,AudioLLM评判器识别显著韵律问题并生成结构化优化指令;优化器(我们的细粒度指令遵循TTS模型)则基于初始话语、目标文本和指令进行引导式表达重合成。为训练优化器,我们构建了Refiner-DB,一个包含4.2万个示例的AudioLLM标注数据集,带有词级韵律弱监督。对经诊断的低质量话语的人工评估显示,LoopTTS可检测感知显著错误并通过优化器校正,在修复质量上优于原始生成音频和实用开环重生成基线;优化器在针对性韵律修改的重音和停顿控制上也展现出更强的指令遵循能力。

英文摘要:

Current TTS systems typically rely on open-loop, single-pass generation and can produce sporadic local prosodic defects, such as misplaced stress, unnatural pauses, or flattened intonation, that utterance-level metrics often fail to expose. We present LoopTTS, a judge-guided Filter-Judge-Refiner framework for recovering low-quality TTS outputs diagnosed by an AudioLLM. Given an initial utterance from a base TTS model, an AudioLLM Judge identifies salient prosodic issues and generates structured refine instructions; a Refiner, our fine-grained instruction-following TTS model, then performs guided expressive re-synthesis conditioned on the initial utterance, target text, and instruction. To train the Refiner, we construct Refiner-DB, a 42K-example AudioLLM-annotated dataset with word-level prosodic weak supervision. Human evaluation on diagnosed low-quality utterances shows that LoopTTS can detect perceptually salient errors and correct them with the Refiner, outperforming raw generated audio and practical open-loop re-generation baselines in recovery quality. The Refiner also demonstrates stronger instruction-following ability for stress and pause control in targeted prosody modification.

补充信息

↑