arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在自动语音识别中从未标记测试数据学习新词

Learning New Words from Unlabeled Test Data in Automatic Speech Recognition

Mengqi Wang, Mark A. Hasegawa-Johnson, Haolong Zheng, Chang D. Yoo

arXiv 2609.28877首次发表:更新:

发表机构

University of Illinois at Urbana-Champaign; Korea Advanced Institute of Science and Technology(伊利诺伊大学厄巴纳-香槟分校; 韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种在测试时从未标记数据学习新词上下文表示与拼写的ASR方法,利用CTC声学模型、语言模型及适应模块,通过KLD优化,在LibriSpeech和构音障碍数据上分别降低OOV字符错误率14.97%和6.67%。

AI 中文摘要

新词每天都在被创造。人类听众可以通过清晰听到一次新词并从句子上下文中推断其用法来学习该词。本文提出赋予自动语音识别(ASR)类似的能力,即在测试时从未标记的测试数据中学习新词的上下文表示和拼写。一个冻结的CTC声学模型提供拼写,一个冻结的语言模型为词汇外(OOV)词检测提供上下文证据,一个适应模块通过学习词汇标记表示(基于CTC生成候选的分布)来扩展词汇表。每个标记的拼写模型通过最小化Kullback-Leibler散度(KLD)目标进行优化。我们证明了CTC加权的语言模型对数似然比可以解释为未知正确ASR与无监督学习ASR之间的KLD,并且利用Pinsker界,KLD的平方根可以解释为未知词的真实拼写与估计拼写之间的总变差距离的上界。实验表明,相对于相应的重打分系统,在LibriSpeech上重复出现的OOV词的相对OOV字符错误率降低了高达14.97%,在构音障碍语音可访问项目数据上降低了6.67%。

英文摘要

New words are invented every day. A human listener can learn a new word by hearing it clearly once and inferring its usage from sentence context. This paper proposes granting ASR a similar ability to learn the contextual representations and spellings of new words from unlabeled test data at test time. A frozen CTC acoustic model provides spellings, a frozen language model provides contextual evidence for out-of-vocabulary (OOV) word detection, and an adaptation module expands the vocabulary by learning the lexical token representations with distributions over CTC-generated candidates. The spelling model of each token is optimized by minimizing a Kullback-Leibler divergence (KLD) objective. We demonstrate that the CTC-weighted language model log likelihood ratio can be interpreted as the KLD between the unknown correct ASR and the unsupervised learned ASR, and that, using a Pinsker bound, the square root of KLD can be interpreted as an upper bound on the total variation distance between the true and estimated spelling of the unknown word. Experiments show relative OOV character-error-rate reductions of up to 14.97% on LibriSpeech and 6.67% on dysarthric Speech Accessibility Project data for recurring OOV words, relative to the corresponding rescoring system.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑