arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DuplexSpeechBench-文档基础:语音代理中文档基础与幻觉的基准测试

DuplexSpeechBench-Document Grounding: Benchmarking Document Grounding and Hallucinations in Voice Agents

Puneet Mathur, Nedim Lipka, Zeyu Jin, Dinesh Manocha

arXiv 2610.00316首次发表:更新:

发表机构

Adobe Research; University of Maryland College Park(Adobe研究院; 马里兰大学学院公园分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出DSB-DG基准,评估语音代理在文档基础中的准确性、幻觉与延迟,发现级联系统表现最佳,开放权重系统存在上下文容量崩溃,基础保真度随负载下降。

AI 中文摘要

语音代理能够实现低延迟、自然的交互,然而其在外部文档中忠实基础响应的能力仍未得到充分探索。我们引入了DuplexSpeechBench-文档基础(DSB-DG),这是一个用于评估语音代理在五个专业领域中文档基础能力的基准。DSB-DG针对三种失败模式:上下文饱和,衡量在文档长度增加情况下的基础能力;基础衰减,衡量在多轮对话中文档事实的保持能力;以及主动基础,评估上下文重新注入是否能缓解对话漂移。该基准包含来自50份文档的1,636个对抗性验证的问答对,覆盖五个专业领域,并支持对基础准确性、幻觉和响应延迟的完全自动评估。在涵盖级联式、专有全双工和实时、以及开放权重语音到语音架构的系统中,我们发现有效基础能力存在显著差异。虽然级联流水线(ASR-LLM-TTS)实现了最高的基础准确性,但Gemini-Live和GPT-Realtime紧随其后。开放权重系统表现出不同的失败模式,最显著的是突然的上下文容量崩溃和多轮基础衰减。更广泛地说,基础保真度随着上下文和对话负载的增加而降低,失败常常表现为无支撑的生成而非弃权(不执行)。我们表明,上下文基础是可靠全双工语音代理面临的关键未解决挑战。

英文摘要

Voice agents enable low-latency, natural interaction, yet their ability to faithfully ground responses in external documents remains underexplored. We introduce DuplexSpeechBench-Document Grounding (DSB-DG), a benchmark for evaluating document grounding in voice agents across five professional domains. DSB-DG targets three failure modes: Context Saturation, which measures grounding under increasing document length; Grounding Decay, which measures retention of document facts across multi-turn dialogue; and Proactive Grounding, which evaluates whether context re-injection mitigates conversational drift. The benchmark contains 1,636 adversarially verified QA pairs from 50 documents covering five professional domains, and supports fully automatic evaluation of grounding accuracy, hallucination, and response latency. Across systems spanning cascaded, proprietary full-duplex and real-time, and open-weight speech2speech architectures, we find substantial differences in effective grounding capacity. While cascaded pipeline (ASR-LLM-TTS) achieves the highest grounding accuracy, Gemini-Live and GPT-Realtime closely trail behind. Open-weight systems exhibit distinct failure modes, most notably an abrupt context-capacity collapse and multi-turn grounding decay. More broadly, grounding fidelity degrades with context and conversational load, and failures frequently manifest as unsupported generations rather than abstention. We show that contextual grounding as a key unresolved challenge for reliable full-duplex voice agents.

CommentsUnder submission at EACL 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑