VAmoS 续篇:更困难、更逼真的语音代理模拟
VAmoS Part Deux: Harder, More Realistic Voice-Agent Simulation
- Veris AI
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对生产语音代理的多请求、背景语音和客户不耐烦挑战,提出 VAmoS Energy 基准,通过 100 个账单电话和 LLM 验证评估,发现完成率低且背景干扰影响大,强调全通话评估的必要性。
AI中文摘要:
生产环境中的语音代理必须处理多个请求、背景语音以及失去耐心的客户。我们引入了 VAmoS Energy,一个基准测试,它在 100 个关于公用事业账单和支付协助的电话中结合了这些挑战。每个来电者提出两到四个请求。该代理拥有十六个工具,由有状态的 Stripe 账单孪生和 Apache Fineract 贷款引擎支持,在来电者验证成功之前,账户访问被阻止。这些任务使用公共家庭电力数据和基于宾夕法尼亚州住宅计费规则的政策。一个 LLM 作为验证器,根据明确要求检查代理的操作和口头数字。在校准运行中,它与代码验证器在 99.1% 的检查上达成一致。在十四个语音栈和每个任务三次重复中,完成率从 17.3% 到 44.7% 不等。Grok Voice 领先,Gemini 3.8 Live 和 GPT-Live 以大致相同的每次通话成本紧随其后。背景电视将合并完成率从 38.7% 降低到 8.6%。模拟来电者经常接受不正确的结果,因为它听到代理的话但无法检查其操作。这些发现表明,语音代理需要在整个通话中进行评估,包括它们说什么、改变什么以及如何处理竞争性语音。
英文摘要:
Voice agents in production must handle several requests, background speech, and customers who lose patience. We introduce VAmoS Energy, a benchmark that combines these challenges in 100 calls about utility billing and payment assistance. Each caller makes two to four requests. The agent has sixteen tools backed by a stateful Stripe billing twin and the Apache Fineract loan engine, with account access blocked until caller verification succeeds. The tasks use public household electricity data and a policy based on Pennsylvania's residential billing rules. An LLM-as-a-verifier checks the agent's actions and spoken figures against explicit requirements. On a calibration run, it agrees with a code verifier on 99.1% of checks. Across fourteen voice stacks and three repeats per task, completion ranges from 17.3% to 44.7%. Grok Voice leads, and Gemini 3.8 Live and GPT-Live follow at about the same cost per call. Background television reduces pooled completion from 38.7% to 8.6%. The simulated caller often accepts an incorrect result because it hears the agent's words but cannot inspect its actions. These findings show why voice agents need evaluation across the whole call, including what they say, what they change, and how they handle competing speech.