arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MSI-Bench:评估协作型AI智能体的多说话人语音交互

MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents

Chenxu Xiong, Dongming Shen, Yuzhi Tang, Wentao Ma, Mu Li, Alex Smola

arXiv 2609.24812首次发表:更新:

发表机构

Boson AI(Boson AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多说话人语音交互场景,提出MSI-Bench基准,含1152个中英测试用例,评估记忆、指令遵循与推理能力,揭示感知与决策瓶颈,为语音智能体指明改进方向。

AI 中文摘要

语音为AI智能体提供了一种自然且即时的交互界面。许多语音智能体可能发挥作用的场景,包括会议、家庭和协作工作,本质上都是多说话人的。支持这些场景带来了在一对一交互中基本不存在的新挑战。我们引入了多说话人交互基准(MSI-Bench),用于评估多说话人语音交互。每个测试用例都是一个短小的多方多轮音频场景,包含参与者上下文、预期工具调用和原子评分细则。该基准针对三个能力族:多说话人记忆、多说话人指令遵循和多说话人推理。它包含1,152个测试用例,在普通话和英语之间均匀分配(各576个)。在每个分割上最强的配置仅在66.8%的英语用例和54.5%的普通话用例中通过了所有评分细则,而最强的开放权重配置分别在34.0%和19.3%的用例中通过。失败分析将感知与推理区分开来:开放权重模型受限于多说话人音频前端,而前沿系统在干净转录文本上仍然无法进行说话人范围的决策——并且各类模型在无人对其说话时也常常做出响应。这些结果将说话人接地感知、说话人范围决策和对话克制确定为未来语音智能体的具体改进目标。

英文摘要

Voice provides a natural and immediate interface for AI agents. Many settings in which voice agents could be useful, including meetings, households, and collaborative work, are inherently multi-speaker. Supporting these settings introduces challenges that are largely absent from one-on-one interaction. We introduce the Multi-Speaker Interaction Benchmark (MSI-Bench) for evaluating multi-speaker voice interaction. Each test case is a short multi-party multi-turn audio scene with participant context, expected tool calls, and atomic rubrics. The benchmark targets three capability families: multi-speaker memory, multi-speaker instruction following, and multi-speaker reasoning. It comprises 1,152 test cases, evenly split between Mandarin Chinese and English (576 each). The strongest configuration on each split passes all rubrics on only 66.8% of English and 54.5% of Mandarin cases, and the strongest open-weight configuration on 34.0% and 19.3%. Failure analysis separates perception from reasoning: open-weight models are bottlenecked by the multi-speaker audio front-end, while frontier systems still fail speaker-scoped decision making on clean transcripts---and models across the board often respond when no one has addressed them. These results identify speaker-grounded perception, speaker-scoped decision making, and conversational restraint as concrete targets for future voice agents.

Comments23 pages, 6 figures, 5 tables. Dataset: https://huggingface.co/datasets/M2cha4l1124/MSI-Bench ; Code: https://github.com/boson-ai/MSI-Bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑