DuplexWorld:语音智能体能帮你度过一天吗?
DuplexWorld: Can voice agents help you get through the day?
- Centific Global Solutions Inc.(森蒂菲克全球解决方案公司)
- University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
DuplexWorld针对现有语音智能体评估基准的不足,构建含六大领域的156个场景开展评估,发现现有最优语音智能体在多维度仍有较大改进空间,并分析了相关性能与失败模式。
AI中文摘要:
语音转语音(S2S)语音智能体因相比文本的会话模态更便捷,正越来越多地被企业用于客户服务,并作为消费者的日常陪伴。然而,现有基准未能沿真正重要的维度全面评估语音智能体,且这些基准被设计为智能体工具调用针对数据库的测试。我们认为,现有基准未充分考虑日常活动带来的会话多样性,也从未测试智能体在超出数据库操作的任务中能提供多忠实的协助。为解决这一问题,DuplexWorld 引入了六个语音智能体特别有用的领域:银行、保险、旅行、医疗、物流和路径规划(Pathfinding)。智能体在156个场景(350多个小时的对话)中的11种不同类型会话上接受评估,每种场景都在不同程度上测试会话和分析能力。通过包含智能体、会话和语音自然度指标的广泛评估,我们表明,即使是最好的语音智能体在所有三个维度上都有很大的改进空间(Pass@1:0.490,轮次交替:0.653,DNSMOS:3.378)。我们对智能体与会话性能、领域和会话类型的性能、失败模式(针对路径规划会话探索了探索与利用的视角)以及所有六个领域中语音智能体的可靠性进行了广泛分析。
英文摘要:
Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing benchmarks fail to holistically evaluate voice agents along axes that really matter and are shaped as tests of agentic tool calling against a database. We believe they fail to adequately account for the diversity of conversational dialogue that mundane activities introduce and further, never test how faithfully an agent can assist on tasks that move beyond database manipulation. To tackle this DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding. Agents are evaluated on eleven different types of conversations across 156 scenarios (350+ hours of conversation), each testing conversational and analytical capability to varying degrees. Through extensive evaluation comprising agentic, conversational and speech-naturalness metrics, we show that even the best voice agents leave substantial room for improvement on all 3 axes (Pass@1: 0.490, turn-taking: 0.653, DNSMOS: 3.378). We perform extensive analysis on agentic v conversational performance, world- and conversation type-wise performance, failure modes exploring the explore v exploit lens for Pathfinding conversations and voice agent reliability over all six worlds.