发表机构
The University of Texas at Austin; National Taiwan University; Massachusetts Institute of Technology; Carnegie Mellon University; Meta AI(德克萨斯大学奥斯汀分校; 国立台湾大学; 麻省理工学院; 卡内基梅隆大学; Meta AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有基准忽视多方对话场景的问题,提出首个多方对话基准MP-Bench,从轮次转换和反应恰当性评估语音智能体,发现实时智能体表现接近随机,揭示重大挑战。
AI 中文摘要
对话式语音智能体已取得显著进展,通过级联和端到端两种架构提供了日益自然的人机交互体验。然而,尽管近期基准测试广泛评估了双人交互和被动音频理解,它们大多忽略了一个普遍存在的现实场景:多方对话。在这些场景中评估智能体从根本上比双人交互更具挑战性,因为对话复杂性呈指数级增长。为使语音智能体无缝融入人类群体动态,它们不仅需要生成上下文恰当的反应,还必须展现出对开放轮次转换的细致理解。为填补这一空白,我们引入了多方基准(MP-Bench),这是首个专门设计用于在多方语境中客观评估对话式语音系统作为积极参与者的基准。MP-Bench 从两个主要维度评估智能体行为:轮次转换意识和反应恰当性。此外,我们纳入基于理解的问答任务作为补充评估。通过对 12 个语音智能体进行基准测试,我们发现实时语音智能体在多方理解任务上的得分不高于 22%,在多方轮次转换上接近随机水平,这揭示了实时语音智能体在多方场景下面临的公开挑战。
英文摘要
Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-end architectures. However, while recent benchmarks extensively evaluate dyadic interactions and passive audio comprehension, they largely overlook a prevalent real-world scenario: multi-party conversations. Evaluating agents in these settings is fundamentally more challenging than in dyadic interactions due to the exponentially greater conversational complexity. For voice agents to integrate seamlessly into human group dynamics, they must not only generate contextually appropriate responses but also demonstrate a nuanced understanding of open turn-taking. To address this gap, we introduce Multiparty Bench (MP-Bench), the first benchmark specifically designed to objectively evaluate conversational speech systems as active participants within multi-party contexts. MP-Bench assesses agent behavior along two primary dimensions: turn-taking awareness and response appropriateness. Additionally, we incorporate comprehension-based question-answering tasks as a complementary evaluation. By benchmarking 12 voice agents, we find that real-time voice agents stay at or below 22% on multiparty comprehension and remain near chance on multiparty turn-taking, exposing an open challenge for real-time voice agents under multiparty scenario.
CommentsAccepted to EMNLP 2026 Findings