发表机构
Veris AI(维瑞斯人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出 VAmoS Bench 语音智能体仿真基准,针对现有基准无法评估语音智能体独立处理电话呼叫能力的空白,在金融服务场景中端到端评估完整语音智能体系统,支持动态排行榜。
AI 中文摘要
生产环境中的语音智能体涵盖级联、端到端语音转换及混合架构。现有语音智能体基准通常仅评估组件质量与对话属性,如词错误率、延迟、自然度和轮次切换,较少衡量智能体能否独立正确处理电话呼叫。联络中心将此称为“ containment”(问题解决率),即自动化系统无需转人工即可解决的电话呼叫占比,部分呼叫的正确结果为拒绝或转接。为填补这一空白,我们推出 VAmoS Bench(Voice Agent Simulation Bench,语音智能体仿真基准),该基准在有状态的客户支持任务中端到端评估完整语音智能体系统。智能体名为 Riley,是虚构银行的信用卡支持代表,可执行冻结、取消、补办或激活卡片操作。100 个场景各提供一个带有私人目标的模拟呼叫者及已初始化的 PostgreSQL 后端,平台利用每个场景生成并激活独立仿真环境,呼叫者通过音频与 Riley 交互,约三分之一场景施加对抗压力。智能体可使用 5 个工具对后端执行真实 SQL,每个场景还定义了二元断言,评分器依据呼叫者与智能体的完整对话轨迹、智能体操作(含工具调用、参数及返回行)评估断言,可检测智能体声称修改卡片却未更新数据库,或虽完成正确数据库变更却泄露受保护信息的情况。该基准首个版本聚焦金融服务,其评估协议支持动态排行榜:可在同一版本上评估其他语音智能体,后续版本可扩展任务与场景。
英文摘要
Production voice agents span cascaded, speech-to-speech, and hybrid architectures. Voice-agent benchmarks typically measure component quality and conversational properties such as word error rate, latency, naturalness, and turn-taking. Fewer measure whether the agent handled a phone call correctly on its own. Contact centers refer to this as ``containment'': the share of phone calls the automated system resolves without handing off to a human. On some phone calls the right outcome is refusal or a redirect. To address this gap, we introduce VAmoS Bench, the Voice Agent Simulation Bench. It measures complete voice-agent systems end to end in a stateful customer-support task. The agent is Riley, a credit-card support representative for a fictional bank who can freeze, cancel, replace, or activate a card. Each of 100 scenarios supplies a simulated caller with a private goal and a seeded PostgreSQL backend. The platform uses each scenario to populate and activate an isolated simulation in which the caller reaches Riley over audio; roughly one-third apply adversarial pressure. The agent can use five tools that execute real SQL against the backend. Each scenario also defines binary assertions. A grader evaluates them against the complete trace of what the caller and agent said and what the agent did, including tool invocations, arguments, and returned rows. This catches an agent that claims to have changed a card without updating the database, as well as one that makes the right database change while disclosing protected information. This first benchmark version focuses on financial services. Its evaluation protocol supports an evolving leaderboard: additional voice agents can be evaluated on the same version, while later versions can expand the tasks and scenarios.
Comments12 pages, 3 figures, 2 tables. Agent implementations: https://github.com/veris-ai/riley-agent