发表机构
Institute of Science Tokyo; Hokkaido University(科学东京大学; 北海道大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DuplexGen框架通过解耦对话的内容、时序与声学,生成更贴近真实对话的合成语音,构建了带多类标注的医患对话语料库,性能优于传统拼接式合成方法。
AI 中文摘要
合成对话语音已成为开发和评估对话语音系统的重要资源。然而,现有的对话合成流程通常先生成对话内容,再通过手工标记或时序规则插入打断、重叠和反馈通道,使得对话时序是预设的而非交互驱动的。我们提出了DuplexGen,一种明确解耦内容、时序与声学的对话合成框架:首先,大语言模型(LLM)生成对话脚本;随后,两个全双工对话模型在实时相互监听的同时执行该脚本,这使得对话时序能自然呈现,同时保留脚本内容;最后,高保真文本到语音(TTS)模型在不改变时序的情况下重新渲染交互。为展示该框架,我们构建了带构建时标注的医患对话语音语料库,标注内容包括单词时间戳、说话人活动、重叠区域和交互事件。实验结果表明,与传统基于拼接的合成方法相比,所提框架生成的对话动态更接近真实对话。
英文摘要
Synthetic conversational speech has become an important resource for developing and evaluating conversational speech systems. However, existing dialogue synthesis pipelines typically generate dialogue content first and then insert interruptions, overlap, and backchannels using handcrafted markers or timing rules, making conversational timing prescribed rather than interaction-driven. We present Agentic-DuplexGen, a dialogue synthesis framework that explicitly decouples content, timing, and acoustics. An LLM first generates the dialogue script, and then two full-duplex conversational models perform the script while listening to each other in real time. This allows conversational timing to emerge naturally while preserving the scripted content. Finally, a high-fidelity text-to-speech model re-renders the interaction without altering its timing. As a demonstration of the proposed framework, we construct a patient--clinician conversational speech corpus with construction-time annotations, including word timestamps, speaker activity, overlap regions, and interaction events. Experimental results show that the proposed framework produces conversational dynamics closer to real dialogue than conventional stitching-based synthesis.