发表机构
AveraLabs(艾维拉实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ConversationalVoice流水线,通过分离、重建和扩展将真实对话转化为全双工训练数据,保持高说话人相似度与语音质量,为后续模型训练提供基础。
AI 中文摘要
全双工语音模型需要保留轮流说话、重叠、打断和反馈行为(backchannel)的训练数据,然而在嘈杂的真实世界录音中,这些信号在不同说话人之间相互纠缠。我们提出 ConversationalVoice,一个将真实双说话人片段转换为三种互补训练数据产物的流水线。(1)分离(Separation)恢复具有稳定说话人分配的说话人专属音轨、规范转录文本以及自然观察到的交互时序。(2)重建(Reconstruction)从固定源转录文本生成匹配语音,重建源说话顺序、停顿和重叠,并添加词级对齐和发音指令。(3)扩展(Expansion)在源上下文、说话人和观察到的交互模式约束下生成新对话。自动说话人验证指标在各阶段保持强劲,同说话人相似度为0.983-0.991,正判别裕度为0.199-0.209。预测语音质量(NISQA MOS)在分离阶段为3.56,重建阶段为4.41,扩展阶段为4.61。基于Gemini的自动评估器为扩展赋予上下文连贯性平均分4.94/5和对话自然度平均分4.80/5。扩展和重建表现出大致相似的交互特征;扩展的轮流、重叠事件、反馈和打断率分别低4.6%、8.0%、13.2%和16.0%。我们仅评估数据属性;全双工模型训练的下游收益留待未来工作。
英文摘要
Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordings. We present Conversational Voice, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts. (1) Separation recovers speaker-specific tracks with stable speaker assignments, a canonical transcript, and naturally observed interaction timing. (2) Reconstruction generates speech in matched voices from a fixed source transcript, reconstructs the source turn order, pauses, and overlaps, and adds word-level alignment and delivery instructions. (3) Expansion generates new dialogue constrained by the source context, speakers, and observed interaction pattern. Automatic speaker-verification metrics remain strong across stages, with same-speaker similarity of 0.983-0.991 and positive discrimination margins of 0.199-0.209. Predicted speech quality (NISQA MOS) is 3.56 for separation, 4.41 for reconstruction, and 4.61 for expansion. A Gemini-based automatic evaluator assigns expansion mean scores of 4.94/5 for contextual coherence and 4.80/5 for dialogue naturalness. Expansion and reconstruction exhibit broadly similar interaction profiles; expansion's turn, overlap-event, backchannel, and interruption rates are 4.6%, 8.0%, 13.2%, and 16.0% lower, respectively. We evaluate data properties only; downstream gains in full-duplex model training remain for future work.
Comments12 pages, 3 figures, 1 table, preprint