发表机构
Northwestern Polytechnical University(西北工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该论文提出一个基于批评驱动迭代细化的智能体框架,用于可控多说话人对话语音合成,通过语句级和场景级批评路由至编辑或重合成,在双语基准上优于直接模型和仅重生成方案。
AI 中文摘要
多说话人对话语音合成需要自然的语音生成、一致的说话人身份、连贯的跨轮次转换,以及对情感、语速和响度等表达属性的细粒度控制。这些要求难以通过一次性生成可靠地满足,尤其是在长篇幅对话中。我们提出一个可控的多说话人对话语音合成框架,将合成过程表述为基于批评的迭代细化。其语音主干ControlEdit-TTS统一了指令跟随合成和自然语言引导的属性编辑,使得无需完全重新生成即可纠正表达错误。该框架进一步执行层级化的语句级和场景级批评,将检测到的问题路由至编辑、重新合成或时序调整。在双语中英对话基准上的实验表明,该框架在语句级指令跟随方面有所改进,在对话级偏好上优于直接对话模型和智能体基线,并且与仅重新生成的替代方案相比,在保持说话人身份的同时实现了更有效的细化。消融研究进一步证实了场景级批评和基于编辑的纠正的益处。
英文摘要
Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to satisfy reliably with one-shot generation, especially in long-form dialogue. We propose a controllable multi-speaker dialogue TTS framework that formulates synthesis as critique-driven iterative refinement. Its speech backbone, ControlEdit-TTS, unifies instruction-following synthesis and natural-language-guided attribute editing, enabling correction of expressive errors without full regeneration. The framework further performs hierarchical utterance-level and scene-level critique, routing detected issues to editing, resynthesis, or timing adjustment. Experiments on a bilingual Chinese--English dialogue benchmark show improved utterance-level instruction following, better dialogue-level preference than direct dialogue models and agentic baselines, and more effective refinement than regeneration-only alternatives while preserving speaker identity. Ablations further confirm the benefits of scene-level critique and edit-based correction.