arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ConversationalVoice:通过源忠实重建与对话扩展从真实对话生成全双工语音数据

ConversationalVoice: Full-Duplex Speech Data from Real Conversations through Source-Faithful Reconstruction and Conversation-Grounded Expansion

Richard Yucheng He, Baodong Cao, Chen Xu, Yihang Liu, Tairan Chen

arXiv 2609.08147首次发表:更新:

发表机构

AveraLabs(艾维拉实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ConversationalVoice流水线,通过分离、重建和扩展将真实对话转化为全双工训练数据,保持高说话人相似度与语音质量,为后续模型训练提供基础。

AI 中文摘要

全双工语音模型需要保留轮流说话、重叠、打断和反馈行为(backchannel)的训练数据,然而在嘈杂的真实世界录音中,这些信号在不同说话人之间相互纠缠。我们提出 ConversationalVoice,一个将真实双说话人片段转换为三种互补训练数据产物的流水线。(1)分离(Separation)恢复具有稳定说话人分配的说话人专属音轨、规范转录文本以及自然观察到的交互时序。(2)重建(Reconstruction)从固定源转录文本生成匹配语音,重建源说话顺序、停顿和重叠,并添加词级对齐和发音指令。(3)扩展(Expansion)在源上下文、说话人和观察到的交互模式约束下生成新对话。自动说话人验证指标在各阶段保持强劲,同说话人相似度为0.983-0.991,正判别裕度为0.199-0.209。预测语音质量(NISQA MOS)在分离阶段为3.56,重建阶段为4.41,扩展阶段为4.61。基于Gemini的自动评估器为扩展赋予上下文连贯性平均分4.94/5和对话自然度平均分4.80/5。扩展和重建表现出大致相似的交互特征;扩展的轮流、重叠事件、反馈和打断率分别低4.6%、8.0%、13.2%和16.0%。我们仅评估数据属性;全双工模型训练的下游收益留待未来工作。

英文摘要

Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordings. We present Conversational Voice, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts. (1) Separation recovers speaker-specific tracks with stable speaker assignments, a canonical transcript, and naturally observed interaction timing. (2) Reconstruction generates speech in matched voices from a fixed source transcript, reconstructs the source turn order, pauses, and overlaps, and adds word-level alignment and delivery instructions. (3) Expansion generates new dialogue constrained by the source context, speakers, and observed interaction pattern. Automatic speaker-verification metrics remain strong across stages, with same-speaker similarity of 0.983-0.991 and positive discrimination margins of 0.199-0.209. Predicted speech quality (NISQA MOS) is 3.56 for separation, 4.41 for reconstruction, and 4.61 for expansion. A Gemini-based automatic evaluator assigns expansion mean scores of 4.94/5 for contextual coherence and 4.80/5 for dialogue naturalness. Expansion and reconstruction exhibit broadly similar interaction profiles; expansion's turn, overlap-event, backchannel, and interruption rates are 4.6%, 8.0%, 13.2%, and 16.0% lower, respectively. We evaluate data properties only; downstream gains in full-duplex model training remain for future work.

Comments12 pages, 3 figures, 1 table, preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑