arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

$τ$-Elicitation:语音代理中多轮实体提取的基准测试

$τ$-Elicitation: Benchmarking multi-turn entity extraction in voice agents

Soham Ray, Victor Barres

arXiv 2609.13602首次发表:更新:

发表机构

Sierra; Mercor(Sierra; Mercor)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出$τ$-Elicitation语音基准测试,发现语音代理在精确实体收集上成功率低,策略选择与错误恢复是核心瓶颈。

AI 中文摘要

语音代理通常需要精确收集姓名、地址、标识符、日期和时间,然而端到端基准测试掩盖了捕获失败的位置。我们引入了$τ$-Elicitation,一个包含200个任务的语音基准测试,涵盖10种实体类型、受控难度、呼叫者真实性和三种环境。一个匹配的文本代理通过了所有任务,但四种语音配置的稳健精确成功率仅在0.14到0.41之间。代理增加了对困难和不熟悉实体的验证,有时也对错误捕获进行验证,但不会针对其最弱的呼叫者语音进行验证;只有24%到37%的已验证错误得到修复。一个规定拼写、回读、纠正和确认的脚手架将稳健Pass$^3$提高了14到31个百分点,但每次呼叫耗时增加21到28秒。拼写变体和重启等真实因素并未显著影响精确成功率;发音错误增加了修复工作量。这些结果确定了策略选择和成功恢复是精确口语实体收集中的核心瓶颈。

英文摘要

Voice agents often need to collect names, addresses, identifiers, dates, and times exactly, yet end-to-end benchmarks obscure where capture fails. We introduce $τ$-Elicitation, a 200-task voice benchmark spanning 10 entity types, controlled difficulty, caller realisms, and three environments. A matched text agent passes all tasks, but four voice configurations achieve robust exact success from 0.14 to 0.41. Agents increase verification for hard and unfamiliar entities and sometimes for incorrect captures, but not for their weakest caller voice; only 24 to 37 percent of verified errors are repaired. A scaffold that prescribes spelling, read-back, correction, and confirmation raises robust Pass$^3$ by 14 to 31 points, at a cost of 21 to 28 seconds per call. Realisms such as spelling variations and restarts do not detectably affect exact success; mispronunciation increases repair effort. These results identify strategy selection and successful recovery as the central bottlenecks in exact spoken entity collection.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑