arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Talk2Agent:文本代理的语音接口基准评测

Talk2Agent: Benchmarking Voice Interfaces for Text Agents

Terumi Chiba, Guangzhi Sun, Zheqi Yuan, Chao Zhang

arXiv 2609.38867首次发表:更新:

发表机构

Tsinghua University; University of Cambridge(清华大学; 剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对语音接口在LLM计算机使用代理中的转录错误问题,提出Talk2Agent基准及免执行评估框架,衡量任务信息保留度,并在WildClawBench上显著提升相关性。

AI 中文摘要

大型语言模型(LLM)计算机使用代理通常使用清晰的书面指令进行评估,尽管语音正日益成为与此类系统交互的流行接口。语音输入引入了额外的失败点:转录错误可能在代理开始推理之前改变任务关键实体、约束或目标,而传统的自动语音识别(ASR)指标并不能直接衡量成功执行所需的信息是否被保留。我们引入了Talk2Agent,一个用于评估语音接口如何有效地将人类口头指令传达给基于LLM的计算机使用代理的基准。Talk2Agent构建了来自WildClawBench和OSWorld任务的人类语音版本,并评估了一系列语音接口,包括专用ASR模型、支持音频的LLM、上下文偏置和基于LLM的本体修复。由于重复执行长时程计算机使用任务成本高昂且具有随机性,我们进一步提出了一种免执行、任务条件化的评估框架,该框架将原始任务评分器投影到可提示寻址的意图上,并衡量语音接口后保留了多少与任务相关的信息。在WildClawBench上,Talk2Agent的免执行原生投影提供了一种实用的、基于执行的语音接口质量度量,与下游任务完成度相关,并在32小时的真实人类语音上将Pearson相关系数相较于词错误率(WER)/字符错误率(CER)提高了0.246。

英文摘要

Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular interface for interacting with such systems. Speech input introduces an additional failure point: transcription errors can alter task-critical entities, constraints, or targets before the agent begins reasoning, while conventional ASR metrics do not directly measure whether the information required for successful execution has been preserved. We introduce Talk2Agent, a benchmark for evaluating how effectively voice interfaces convey human-spoken instructions to LLM-based computer-use agents. Talk2Agent builds human-spoken versions of tasks from WildClawBench and OSWorld and evaluates a range of voice interfaces, including dedicated ASR models, audio-capable LLMs, contextual biasing, and LLM-based ontology repair. Because repeatedly executing long-horizon computer-use tasks is costly and stochastic, we further propose an execution-free, task-conditioned evaluation framework that projects the original task grader onto prompt-addressable intentions and measures how much task-relevant information is retained after the voice interface. On WildClawBench, Talk2Agent's execution-free native projection provides a practical, execution-grounded measure of voice-interface quality, correlating with downstream task completion and improving Pearson correlation by 0.246 over WER/CER on 32 hours of real human speech.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑