arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30798cs.AI

评估实时语音智能体:从组件质量到基于事实的结果

Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes

Shivam Negi, Arpit Rawat, Rashi Jain

首次发表
浏览论文内容

中文总结 AI 辅助

针对实时语音智能体评估文献分散的问题,基于38个来源提出三项证据主张,指出评估转向基于事实的结果,并提出TRG报告标准,以时序、恢复和状态验证结果综合表征智能体。

中文摘要 AI 辅助

实时语音智能体已从研究原型走向生产部署,然而描述它们的文献分散在三个很少相互引用的社区中:语音基础建模、话轮转换心理语言学和智能体评估。架构论文报告延迟,话轮转换论文报告预测准确率,智能体基准报告任务成功率,因此没有一个单一数字能描述已部署的智能体是否真正优秀。我们通过三项基于证据的主张来弥补这一空白,每项主张都可追溯到由38个主要来源组成的语料库,该语料库按应用中心分类法组织为六个类别。首先,架构选择是部署约束而非定论:一份2026年企业教程报告称,尚无完全可自托管的端到端系统满足生产约束,而分块级联独立达到最先进的双工行为,表明双工行为可与双工架构分离。其次,评估已从组件质量决定性地转向基于事实的结果,最近的基准验证后端状态,而非相信智能体声称已执行的操作。第三,大多数模型和基准中的二元假设正在瓦解:多方话轮转换和多说话人推理基准表明,决定何时不发言以及推理谁可被告知什么,是双参与框架无法衡量的头等能力。对于每个来源,我们陈述其针对的问题、机制和报告的证据,以及搜索策略、纳入标准和验证步骤,该验证步骤捕获了流通中一个错误归属的arXiv标识符。我们提出TRG(时序-恢复-基于事实),一种报告标准,通过时序、中断后恢复和状态验证结果共同表征智能体,并为多方部署提供条件性第四轴。

英文摘要

Real-time voice agents have moved from research prototypes to production deployments, yet the literature describing them is fragmented across three communities that rarely cite one another: speech foundation modelling, turn-taking psycholinguistics, and agentic evaluation. Architecture papers report latency, turn-taking papers report prediction accuracy, and agentic benchmarks report task success, so no single number describes whether a deployed agent is actually good. We address that gap with three evidence-based claims, each traceable to a corpus of 38 primary sources organised into an application-centric taxonomy of six categories. First, architecture choice is a deployment constraint rather than a settled verdict: a 2026 enterprise tutorial reports that no fully self-hostable end-to-end system yet meets production constraints, while a chunked cascade independently reaches state-of-the-art duplex behaviour, showing duplex behaviour is separable from duplex architecture. Second, evaluation has shifted decisively from component quality toward grounded outcomes, with recent benchmarks verifying backend state rather than trusting what the agent claims to have done. Third, the dyadic assumption in most models and benchmarks is breaking down: multiparty turn-taking and multi-speaker reasoning benchmarks show that deciding when not to speak, and reasoning about who may be told what, are first-class capabilities two-participant framings cannot measure. For each source we state the problem it targets, its mechanism, and its reported evidence, alongside the search strategy, inclusion criteria, and a verification step that caught a misattributed arXiv identifier in circulation. We propose TRG (Timing-Recovery-Grounded), a reporting standard characterising an agent by timing, post-disruption recovery, and state-verified outcome together, with a conditional fourth axis for multiparty deployments.

发表机构

  • Northeastern University(东北大学)
  • Walmart(沃尔玛)
  • Sinch Mailgun
  • University of Texas at Dallas(得克萨斯大学达拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑