arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29299eess.AS

利用束搜索信息进行端到端自动语音识别(ASR)的置信度估计

Leveraging Beam Search Information for Confidence Estimation in E2E ASR

Yichen Jia, Hugo Van hamme

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出轻量型SR-CEM模块,利用束搜索信息生成令牌级与词级置信度得分,在域内、域外英语及多语言、多架构ASR场景中,其最大校准误差显著低于softmax置信度,提升了ASR输出的可靠性。

中文摘要 AI 辅助

为了估计端到端自动语音识别(ASR)系统的置信度,近期研究提出了结合骨干ASR模型特征的置信度估计模块。然而,现有多数方法依赖于特定架构。本文提出Score-Rank置信度估计模块(SR-CEM),这是一种轻量型模块,利用束搜索信息生成令牌级和词级置信度得分。具体而言,SR-CEM通过结合假设内令牌的得分和排名来构建特征。实验表明,SR-CEM在域内和域外英语数据上均实现了有效的校准:在域内测试集上,令牌级的最大校准误差为4.50%,预期校准误差为0.30%,显著优于softmax置信度的20.04%和1.75%;词级的最大校准误差为8.17%,预期校准误差为0.35%,优于softmax置信度的17.91%和1.67%。此外,研究还验证了其在混合ASR架构、 transducer ASR架构及不同解码策略下,以及荷兰语、带噪语音和对话式语音场景中的鲁棒性。核心发现为,SR-CEM在降低最大校准误差方面效果尤为显著,该误差对ASR输出的可靠下游应用至关重要,同时其具有架构独立性,可在多种评估场景中保持通用性。

英文摘要

To estimate confidence for end-to-end Automatic Speech Recognition (ASR) systems, recent research has proposed Confidence Estimation Modules that incorporate features from the backbone ASR model. Most existing approaches, however, are architecture-dependent. In this paper, we propose the Score-Rank Confidence Estimation Module (SR-CEM), a lightweight module that leverages beam search information to generate token- and word-level confidence scores. Specifically, SR-CEM constructs features by combining the scores and ranks of tokens within a hypothesis. Experiments show that SR-CEM achieves effective calibration on both in-domain and out-of-domain English data. On the in-domain testset, it attains a Maximum Calibration Error of 4.50% and an Expected Calibration Error of 0.30% at the token level, significantly outperforming softmax confidence (20.04% and 1.75%, respectively). At the word level, SR-CEM achieves 8.17% and 0.35%, compared to 17.91% and 1.67% from softmax confidence. Furthermore, we demonstrate its robustness across hybrid and transducer ASR architectures with different decoding strategies, as well as on Dutch, noisy and conversational speech conditions. Our main finding is that SR-CEM is particularly effective in reducing Maximum Calibration Error, which is critical for reliable downstream use of ASR outputs, while maintaining architecture independence and generality across diverse evaluation conditions.

补充信息

↑