arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非所有语音都是意图:用于ASR后误唤醒的自适应自校正推理层

Not All Speech Is Intent: Adaptive Self-Correcting Inference Layer for Post-ASR False Wake-Up

Preeti Saraswat, Divya Neelagiri, Anil Yadav

arXiv 2609.12469首次发表:更新:

发表机构

Samsung Research America(三星美国研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对对话AI中ASR后误唤醒问题,提出ASCIL框架,融合声学、语言、上下文及历史错误信号进行自适应校正,在专有数据集上实现显著错误率降低且延迟极低。

AI 中文摘要

误唤醒激活仍然是对话式AI中一个持续存在的挑战。与设备唤醒词语音上相似的语音可能产生语法有效且语义连贯的ASR转录文本,而助手会错误地执行该文本。大多数现有系统孤立地做出单一意图决策,没有机制从随时间重复出现的错误中学习,也没有通过个性化学习适应个体用户。我们引入了反馈驱动的自适应自校正推理层(ASCIL),这是一个互补的ASR后校正框架,在响应生成之前通过融合声学嵌入、语言线索、设备上下文以及过去错误分类的模式来重新评估唤醒意图。ASCIL将隐式信号(包括犹豫、脱离和沉默)和显式信号(包括取消和重复)解释为自动推断的、可能错误分类的噪声行为指标。这些信号驱动在线模式更新,无需人工标注,而用于离线评估的有意/无意参考标签则由人工标注。它从先前的错误中泛化,在推理时应用校正调整,并与自然语言执行并行持续更新。在一个包含3,667次交互的专有数据集上评估,该数据集具有跨越14种声学和上下文条件的人工标注有意/无意参考标签,ASCIL在基于基线失败构建的会话不相交子集上实现了54.27%的相对错误减少,在阈值0.90的问题标记评估切片上实现了高达24.39%的相对错误减少。这些收益是在提高有意接受率的同时实现的,在报告的基准测试中,中位增加延迟低于60毫秒。

英文摘要

False wake-up activations remain a persistent challenge in conversational AI. Speech phonetically similar to a device's wake word can produce a syntactically valid and semantically coherent ASR transcript that the assistant incorrectly executes. Most existing systems make a single intent decision in isolation, without a mechanism to learn from recurring errors over time or adapt to individual users through personalized learning. We introduce the Feedback-Driven Adaptive Self-Correcting Inference Layer (ASCIL), a complementary post-ASR correction framework that re-evaluates wake-up intent before response generation by fusing acoustic embeddings, linguistic cues, device context, and patterns from past misclassifications. ASCIL interprets implicit signals, including hesitation, disengagement, and silence, and explicit signals, including cancellation and repetition, as automatically inferred, noisy behavioral indicators of potential misclassification. These signals drive online pattern updates without manual annotation, whereas the intentional/unintentional reference labels used for offline evaluation are human-annotated. It generalizes from prior errors, applies corrective adjustments at inference time, and continuously updates in parallel with natural-language execution. Evaluated on a proprietary dataset of 3,667 interactions with human-annotated intentional/unintentional reference labels spanning 14 acoustic and contextual conditions, ASCIL achieves 54.27% relative error reduction on a session-disjoint subset constructed from baseline failures, and up to 24.39% relative error reduction at threshold 0.90 on the issue-tagged evaluation slice. These gains are achieved while improving intentional acceptance rates, with a median added latency below 60 ms in the reported benchmark.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑