arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

权衡存在于标签中:面向轮次感知的流式自动语音识别(ASR)的因果监督

The Trade-off Was in the Labels: Causal Supervision for Turn-Aware Streaming ASR

Bojie Li, Noah Shi

arXiv 2609.04225首次发表:更新:

发表机构

Pine AI; University of Washington(Pine AI; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对流式ASR的轮次检测问题,提出基于Qwen3-ASR-0.6B的LoRA适配器,采用因果监督训练,在部署匹配基准上实现了优于沉默超时方案的轮次检测性能,解决了标签的先知式信息漏洞。

AI 中文摘要

语音智能体必须时刻判断用户是否完成发言;沉默极少能明确指示:呼叫者在读电话号码时会在数字间停顿,一个单词“停止!”即可结束一轮发言,长问题的停顿时长会超过实际轮次间隔。部署的默认方案(语音活动检测器加沉默超时机制)无法区分这些情况,因为轮内停顿通常会超过轮间间隔;区分的关键在于目前的话语是否构成完整语义:这是识别器生成转录文本的依据。我们提出了首个轮次感知流式ASR的开源训练方案与基准:基于Qwen3-ASR-0.6B的小型LoRA适配器,在单GPU上训练仅需数小时,该适配器可转录语音、从语义与沉默中检测轮次结束、处理听写任务并将转录文本与上下文关联。在与部署场景匹配的基准上,它达到了0.97的边界召回率,中位数延迟为0.39秒,每语音分钟误触发0.3次,且在全新测试集上可复现;没有任何沉默超时方案能达到该性能。该方案基于一条原则:每个流式决策标签必须可由决策点之前的输入计算得出。离线语料库违反了该原则,其编码了未来信息;这种“先知式”标签会导致振荡和虚假的召回率-精确率权衡,例如在追加1秒沉默后,“受损”模型的轮次结束召回率从0.10升至1.00,暴露了这一问题。同样的漏洞也出现在上下文处理中:始终匹配的偏置前缀成为被复制的捷径(40%的入侵率),而与音频不符的反事实将该入侵率降至0.8%,同时保留了+28.9个百分点的实体召回率增益。

英文摘要

A voice agent must decide, moment to moment, whether the user has finished; silence rarely settles it: a caller reading a phone number pauses mid-digits, a one-word "Stop!" ends a turn, a long question carries pauses longer than real turn-gaps. A voice-activity detector plus a silence timeout (the deployed default) cannot separate these, because within-turn pauses routinely exceed between-turn gaps; what distinguishes them is whether the words so far form a complete thought: what a recognizer computes to produce a transcript. We present the first open training recipe and benchmark for turn-aware streaming ASR: a small LoRA adapter on Qwen3-ASR-0.6B, trained in hours on one GPU, that transcribes, detects end-of-turn from meaning and silence, handles dictation, and grounds transcription in context. On a deployment-matched benchmark it reaches 0.97 boundary recall at 0.39 s median latency with 0.3 false fires per speech-minute, replicated on a fresh test set; no silence timeout reaches this point. The recipe rests on one principle: every streaming-decision label must be computable from input up to the decision point. Offline corpora violate it, encoding the future; such clairvoyant labels manufactured oscillation and a phantom recall-versus-precision trade-off, exposed when one appended second of silence raised a "broken" model's end-of-turn recall from 0.10 to 1.00. The same leak recurred with context: an always-matching biasing prefix became a copied shortcut (40% intrusion), and counterfactuals disagreeing with the audio cut this to 0.8% while keeping most of a +28.9 pp entity-recall benefit.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑