M3-DuplexBench:面向全双工口语对话模型的多轮、多语言、多领域基准测试集
M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models
浏览论文内容
中文总结 AI 辅助
研究针对全双工口语对话系统多轮公平比较难、现有基准覆盖不足的问题,提出支持英日双语、涵盖两类对话场景的M3-DuplexBench基准,通过多设置实验揭示模型话轮特征与性能差异。
中文摘要 AI 辅助
全双工口语对话系统(FDSDS)可在说话时监听,实现流畅话轮转换、反馈处理、用户插话处理等自然行为。但多轮对话中的公平比较仍是挑战,现有基准对语言和对话领域的覆盖有限。我们提出M3-DuplexBench,这是面向FDSDS的多轮、多语言、多领域基准测试集,支持英语和日语,涵盖日常对话与多轮问答。此外,我们在单轮、仅用户、教师强制全上下文等多种对话上下文设置下评估模型,分析对话历史对模型行为的影响。对近期FDSDS的实验显示,模型具有特定的话轮转换特征,不同语言和领域间存在明显性能差距,对话上下文的影响则各不相同。
英文摘要
Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling. However, fair comparisons in multi-turn conversations remain a challenge. In addition, existing benchmarks provide limited coverage of languages and dialogue domains. We propose M3-DuplexBench, a multi-turn, multilingual, multidomain benchmark for FDSDSs. M3-DuplexBench supports English and Japanese and covers both casual conversation and multi-turn question answering. In addition, we evaluate models under multiple dialogue context settings, including single-turn, user-only, and teacher-forced full-context settings, to analyze how dialogue history affects model behavior. Experiments with recent FDSDSs reveal model-specific turn-taking characteristics, clear performance gaps across languages and domains, and mixed effects of dialogue context.