arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28806eess.AScs.AI

一种用于合成多样化自然全双工对话的生成框架

A Harness for Synthesizing Diverse Naturalistic Full-Duplex Conversations

  • Tencent Americas(腾讯美洲)
  • Wuhan University(武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

Matthew Sun, Vinay Kothapally, Meng Yu, Chao Huang, Hao Zhang, Yixuan Zhang, Steve Yves

AI总结:

提出一种从关系事件列表合成带意图标注的全双工对话语音的流水线,覆盖42种现象,通过LLM生成事件并置于共享时钟,实验表明合成数据可提升全双工话轮管理性能。

AI中文摘要:

全双工对话系统在说话的同时需要监听,必须区分完整话轮结束与话轮内停顿、请求话轮的打断与简短确认或对第三方说话。然而,现有对话语料库对这些事件的控制有限,且对其意图的标注也有限。我们提出了一种从关系事件列表合成带意图标注的双声道对话语音的流水线。大语言模型(LLM)为每个事件生成说话人、文本、对话行为和与先前事件的关联,而不预测绝对时间戳。事件被独立合成,与其源文本对齐,并放置在共享时钟上,从而从渲染信号中测量话轮转换标志,同时静音时长由指定或从话轮转换分布中采样。该流水线覆盖英语和中文中八个类别的42种现象,从作者标注的意图推导帧级系统动作,并通过使用小型多样的先验示例集和请求替代方案(带自报概率)的批量提示来促进多样性。消融实验显示每个目标多样性维度均有提升。在用于获取、保持、释放和不保持对话话轮的四动作标签空间中,仅使用当前和过去音频的语义语音活动检测器在开始说话和开始聆听的F1分数分别达到0.819和0.802。在生成自身响应时,全双工语音模型Moshi在生成语料上微调后,参考话轮的0.85被采用,而微调前为0.44。其预测系统话轮占用的帧级精确率从0.46提升至0.88。在每一步使用参考上下文时,其帧级话轮F1从0.893提升至0.962。这些结果表明,受控合成可以为全双工话轮管理提供可学习且可迁移的监督信号。

英文摘要:

Full-duplex dialogue systems, which listen while speaking, must distinguish a completed turn from a pause within a turn and an interruption that requests a turn from a brief acknowledgment or speech addressed to a third party. Yet existing conversational corpora provide limited control over these events and limited labels for their intent. We present a pipeline for synthesizing intent-labeled, two-channel conversational speech from relational event lists. An LLM authors each event's speaker, text, conversational act, and attachment to an earlier event without predicting absolute timestamps. Events are synthesized independently, aligned with their source text, and placed on a shared clock, so turn-taking landmarks are measured from the rendered signal while silence durations are specified or sampled from turn-taking distributions. The pipeline covers 42 phenomena across eight families in English and Mandarin, derives frame-level system actions from authored intent, and promotes diversity using small, diverse sets of prior examples and batch prompts that request alternatives with self-reported probabilities. Ablations show gains in each targeted diversity dimension. On a four-action label space for taking, holding, releasing, and not holding the conversational floor, a semantic voice-activity detector using only current and past audio reaches start-speaking and start-listening F1 scores of 0.819 and 0.802. When generating its own responses, the full-duplex speech model Moshi takes 0.85 of the reference turns after fine-tuning on the generated corpus, compared with 0.44 before fine-tuning. Its frame-level precision for predicting system-floor occupancy rises from 0.46 to 0.88. With reference context at each step, its frame-level floor F1 rises from 0.893 to 0.962. These results show that controlled synthesis can provide learnable and transferable supervision for full-duplex turn management.

补充信息

↑