arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MTVA-Bench:评估级联语音代理内部的语言模型

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

Pritish Mishra, Ishaan Kumar, Akshat Mandoli, Sudarshan Kamath

arXiv 2609.20152首次发表:更新:

AI 中文总结

MTVA-Bench在级联语音代理的真实条件下评估语言模型,通过模拟来电者和后端,结合确定性检查与双LLM评判,发现模型间总分差距主要源于参数值、动作排序和规则遵守。

AI 中文摘要

通常,大多数语音代理是级联系统,即ASR模型转录来电者的音频,语言模型读取转录内容并决定说什么以及调用哪些后端工具,TTS模型则说出回复。几乎所有决策都发生在语言模型中,但现有评估要么过于宽泛,要么过于狭窄。端到端语音基准测试对整个流程进行评分,因此识别错误和模型错误混合成一个数字。LLM基准测试将模型隔离,但它们不评估使真实电话变得困难的因素,例如转录问题、来电者的语音被分割到多条消息中,以及回复必须遵循指定的语言和脚本。我们引入了多轮语音代理基准测试(MTVA-Bench),它在与级联系统内部相同的条件下评估语言模型。来电者由遵循一组规则的LLM扮演,工具调用由模拟后端回答,该后端响应模型实际发送的参数。该基准测试包含49个代理,涵盖490个经过审查的场景,并支持7种语言。评分是工具调用上的确定性检查与两个LLM评判员的组合,一个评判场景特定规则,另一个在不访问任务的情况下评估对话质量。两个评判员都必须引用转录中的特定消息。任务和对话得分权重相等,因为通话可以完成任务但仍然对来电者不利。在一项七模型研究中,六个模型选择的正确工具彼此相差在6.4分以内,但它们的总分相差24.4分。大部分差距来自参数值、动作排序、规则遵守以及模型在工具调用周围所说的内容。

英文摘要

Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller's audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. Nearly all of the decision making happens in the language model, but existing evaluations measure it either too broadly or too narrowly. End-to-end voice benchmarks score the full pipeline, so recognition errors and model errors mix into a single number. LLM benchmarks isolate the model but they do not evaluate what makes real phone calls hard, such as transcription issues, caller's voice being split across messages and the requirement that replies follow the language and script specified. We introduce the Multi-Turn Voice Agent Benchmark (MTVA-Bench), which evaluates the language model on the same conditions it faces inside a cascaded system. The caller is played by an LLM following a set of rubrics and tool calls are answered by a mock backend which responds to the arguments the model actually sent. The benchmark contains 49 agents working across 490 reviewed scenarios and supports 7 languages. Scoring is a combination of deterministic checks on tool calls with two LLM judges, one that scores scenario specific rules and one that grades conversation quality without access to the task. Both judges must cite specific messages from the transcript. Task and conversation scores are weighted equally, since a call can complete its task and still go badly for the caller. In a seven-model study, six of the models select the correct tool within 6.4 points of one another, but their overall scores span 24.4 points. Most of the gap comes from argument values, action ordering, rule compliance, and what the model says around its tool calls.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑