arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22101cs.LG

超越全新启动:会话语音智能体中流式ASR的有状态推理

Beyond Fresh Starts: Stateful Inference for Streaming ASR in Conversational Voice Agents

Sameep Chattopadhyay, Alexander Erdmann, Mari Ostendorf

首次发表
浏览论文内容

中文总结 AI 辅助

针对会话语音智能体流式ASR因每轮重置状态导致的话语起始性能下降问题,提出两种保留跨轮次话语上下文的状态管理策略,在实验中实现15-21%的相对WER降低。

中文摘要 AI 辅助

现代语音智能体系统依赖于受严格延迟约束运行的流式语音识别模型。本研究表明,由于实时处理的内存限制,这些系统会受到长停顿、反馈通道等会话现象的不利影响。虽然许多智能体流程通过在每轮重置状态来缓解该问题,但此方法会丢弃重要上下文,损害轮次起始时的性能。我们提出两种状态管理策略,保留跨轮次话语的上下文以减少起始错误。在两个最先进的流式模型、两个口语对话基准上开展的实验显示,我们的最优方法在话语起始处取得了15-21%的相对词错误率(WER)降低。

英文摘要

Modern voice-agent systems rely on streaming speech recognition models that operate under stringent latency constraints. This study shows that, due to the limited memory constraints of real-time processing, these systems are adversely impacted by conversational phenomena such as long silences and backchannels. While many agentic pipelines mitigate this by resetting state at each turn, this approach discards vital context and impairs performance at turn onsets. We propose two state-management strategies that preserve cross-utterance context to reduce onset errors. In experiments with two state-of-the-art streaming models on two spoken dialogue benchmarks, our best method yields an average of 15-21% relative WER reduction at utterance onsets.

发表机构

  • University of Washington(华盛顿大学)
  • SRI International(国际 SRI 研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑