MultiTalk:将全双工语音模型扩展到长时、多方、双语对话
MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation
- CUHK MMLab(香港中文大学多媒体实验室)
- CPII under InnoHK(香港物流及供应链管理应用技术研发中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出MultiTalk,通过57.6k小时合成数据训练双语全双工模型,并构建MultiTalkBench基准,显著提升长时多方对话性能。
AI中文摘要:
端到端全双工语音模型已使开源机器对话更接近类人交互,但现有系统在两个相互交织的维度上仍存在局限:长上下文鲁棒性和多方交互。会议、小组课程和社交机器人接待等现实场景要求单一模型在长时间内跟踪、情境化并响应多个说话者。进展受到数据和评估两方面的制约:开放的多方语音语料库规模较小,且并非为编解码器帧级全双工建模而设计;而现有的长音频基准侧重于被动聆听,语音到语音基准则大多为短时和双人对话。我们沿长时程和多方两个轴共同扩展了Moshi范式,涵盖英语和中文。首先,我们发布了57.6k小时的合成训练数据(MultiTalkPT和MultiTalkFT),用于长格式、多方、英汉全双工对话,具有可控的时长、参与者数量、轮流发言、重叠、回馈、打断、称呼对象转换和长距离共指。其次,我们引入了MultiTalkBench,基于真实人类录音构建,用于评估长格式、多方、双语全双工对话。对话平均时长为32.6分钟,并包含长距离实体跟踪、主题连贯性和称呼对象选择的探针。第三,我们训练了一个双语Moshi风格模型,该模型能在长时间内维持连贯的多方英汉对话,并在MultiTalkBench上显著优于开源基线模型,包括Moshi、MiniCPM-o-4.5和Qwen3-Omni-30B-A3B-Instruct。
英文摘要:
End-to-end full-duplex speech models have brought open-source machine conversation closer to human-like interaction, yet existing systems remain limited in two intertwined dimensions: long-context robustness and multi-party interaction. Real-world scenarios such as meetings, group lessons, and social-robot reception require a single model to track, contextualize, and respond to multiple speakers over extended durations. Progress is constrained by both data and evaluation: open multi-party speech corpora remain small and are not designed for codec-frame-level full-duplex modeling, while existing long-audio benchmarks focus on passive listening and speech-to-speech benchmarks are mostly short and dyadic. We extend the Moshi paradigm jointly along the long-horizon and multi-party axes in English and Chinese. First, we release 57.6k hours of synthetic training data ($\href{https://huggingface.co/datasets/MultiTalk/MultiTalkPT}{MultiTalkPT}$ and $\href{https://huggingface.co/datasets/MultiTalk/MultiTalkFT}{MultiTalkFT}$) for long-form, multi-party, English-Chinese full-duplex dialogue, with controllable length, participant count, turn-taking, overlap, backchannels, interruptions, addressee shifts, and long-range coreference. Second, we introduce $\href{https://huggingface.co/datasets/MultiTalk/MultiTalkBench}{MultiTalkBench}$, built from real human recordings, for evaluating long-form, multi-party, bilingual full-duplex dialogue. Conversations average 32.6 minutes and include probes for long-range entity tracking, topic coherence, and addressee selection. Third, we train a bilingual Moshi-style model that sustains coherent multi-party English-Chinese conversations over extended durations and substantially outperforms open-source baselines including Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct on MultiTalkBench.