arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34973cs.AI

APEX-Voice:语音智能体能否通过全双工交互完成专业工作流程

APEX-Voice: Can Voice Agents Complete Professional Workflows Through Full-Duplex Interaction

Puneet Mathur, Dinesh Manocha

首次发表
浏览论文内容

中文总结 AI 辅助

APEX-Voice基准测试通过120个专业工作流程评估五个前沿语音智能体,结果显示无一超过25%的Pass@1,证明流畅对话不等于可靠完成专业工作。

中文摘要 AI 辅助

全双工语音智能体现在可以在语音交互过程中倾听、说话、使用工具并采取行动,但流畅的对话并不能保证正确完成被委托的专业工作流程。我们引入了APEX-Voice,这是一个包含120个交互式专业工作流程的基准测试,涵盖十种工作原型,如表单填写、企业谈判、协调、咨询和面试。每个工作流程都在一个有状态的语音工作台环境中执行,该环境具有特定任务的知识、类型化工具、带金标注的最终工作成果、授权约束以及一个由经过验证的预编译语音实现支持的用户模拟策略。我们同时评估工件字段准确性和端到端工作流程成功率,后者要求正确的终止状态、有效的流程、完成的操作以及有效的最终工件。在五个前沿实时语音智能体——GPT-Live-1、Gemini-3.8-Live、Grok-Voice-Think-2.0、Step-Audio3和GPT-realtime-2.1中,没有一个超过25%的Pass@1,而最佳的Reliable@3仅为10.8%。此外,有状态协调是各系统的主要失败点,而在需要更多知识检索和语音中纠正的工作流程上,成功率进一步下降。总体而言,APEX-Voice是第一个评估语音智能体能否将对话能力转化为可靠专业工作的基准测试。

英文摘要

Full-duplex voice agents can now listen, speak, use tools, and act during spoken interactions, but fluent dialogue does not guarantee correct completion of delegated professional workflows. We introduce APEX-Voice, a benchmark of 120 interactive professional workflows spanning ten work archetypes such as form completion, corporate negotiation, coordination, consulting, and interviewing. Each workflow executes in a stateful Voice Workbench environment with task-specific knowledge, typed tools, gold-annotated final work artifact, authorization constraints, and a user simulation policy backed by validated, pre-compiled speech realizations. We evaluate both artifact field accuracy and end-to-end workflow success, which requires the correct terminal state, valid process, completed actions, and a valid final artifact. Across five frontier real-time voice agents-GPT-Live-1, Gemini-3.8-Live, Grok-Voice-Think-2.0, Step-Audio3, and GPT-realtime-2.1, none exceeds 25% Pass@1, and the best Reliable@3 is only 10.8%. Moreover, stateful coordination is the dominant failure point across systems, while success decreases further on workflows requiring greater knowledge retrieval and mid-speech corrections. Overall, APEX-Voice is the first benchmark for evaluating whether voice agents can translate conversational competence into dependable professional work.

发表机构

  • University of Maryland College Park(马里兰大学学院公园分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑