AI 中文总结
提出Duplex-MPE基准,用于评测全双工对话中语音助手在多方场景下的回答、沉默与停止能力,包含2,000个场景和四个评分指标,实验表明MiniCPM-o 4.5在多数能力上领先。
AI 中文摘要
实时全双工语音模型可以在说话的同时进行聆听,从而实现无需严格轮流边界的自然交互。现有基准评测了轮流说话、打断处理和多方对话,但大多以指定的用户为中心,而非评测助手参与多人共享对话的情形。我们提出Duplex-MPE,用于评估此类助手何时应回答、保持沉默或停止说话。该基准包含2,000个场景,每个场景有三或四名人类说话者和一名助手,每个场景均配对同一请求的显式与隐式指称两种情形。模型接收连续的对话音频,不提供转写文本或预设的轮流边界。四个评分指标分别衡量:新回应发起、回答准确性、沉默保持以及当人类解决请求时停止说话。我们评估了五个开源权重语音系统:MiniCPM-o 4.5、Moshi、FLM-Audio、Voila和Freeze-Omni。MiniCPM-o 4.5在三个评分能力上领先,而其他系统频繁的语音输出可能与不准确的回答或未能保持沉默共存。基于转写文本的Gemini 3.1 Pro参考系统对显式请求的回应频率比隐式请求高出64.3个百分点;配对检验未检测到语音系统在回应率上的显著差异。
英文摘要
Real-time full-duplex speech models can listen while speaking, enabling natural interaction without rigid turn boundaries. Existing benchmarks evaluate turn-taking, interruption handling and multi-round dialogue, but largely centre on a designated user rather than an assistant participating in a shared conversation among several people. We introduce Duplex-MPE to evaluate when such an assistant should answer, remain silent or stop speaking. The benchmark contains 2,000 scenarios with three or four human speakers and one assistant, each paired across explicit and implicit addressing of the same request. Models receive continuous conversation audio without transcripts or supplied turn boundaries. Four scores measure fresh response initiation, answer accuracy, silence preservation and stopping when a human resolves a request. We evaluate five open-weight speech systems: MiniCPM-o 4.5, Moshi, FLM-Audio, Voila and Freeze-Omni. MiniCPM-o 4.5 leads on three scored capabilities, while frequent speech from other systems can coexist with inaccurate answers or failures to remain silent. A transcript-based Gemini 3.1 Pro reference responds 64.3 percentage points more often to explicit than implicit requests; paired tests detect no significant response-rate difference for the speech systems.
Comments27 pages, 3 figures. Project page: https://step-out.github.io/Duplex-MPE-Page/