BaLEEN:利用潜在编码实体进行上下文感知自动语音识别的偏置方法
BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR
浏览论文内容
中文总结 AI 辅助
BaLEEN提出一种基于超网络的轻量级上下文适配方法,利用预训练语言模型编码关键词并注入偏置向量,在不微调ASR模型的情况下,将关键词漏检率降低8.7%,词错误率提升21%,字符错误率提升28%。
中文摘要 AI 辅助
转录领域特定实体和罕见专有名词仍然是自动语音识别(ASR)中的主要挑战。在本文中,我们提出了BaLEEN(利用潜在编码实体进行偏置),一个轻量级的、基于超网络的框架,用于在不微调底层ASR模型的情况下进行动态上下文适应。BaLEEN使用预训练语言模型对可变长度的上下文关键词进行编码,通过Perceiver瓶颈将其压缩为固定序列的潜在向量,并将上下文相关的偏置向量直接注入ASR模型的中间编码器表示中。由于语言模型和骨干ASR模型在训练期间完全保持冻结,BaLEEN作为一个即插即用的适配器,在推理时当上下文偏置被预计算时,零计算开销。我们在一个基于CTC的ASR模型上,使用一个带有注释命名实体和合成语音的维基百科衍生语料库评估了我们的方法。实验结果表明,相对于无偏置基线,BaLEEN在测试集上将关键词漏检率降低了8.7%,同时整体词错误率提高了21%,字符错误率提高了28%。
英文摘要
Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propose BaLEEN (Biasing with Latent Encoded Entities), a lightweight, hypernetwork-based framework for dynamic contextual adaptation without fine-tuning the underlying ASR model. BaLEEN encodes variable-length contextual keywords using a pretrained language model, compresses them into a fixed sequence of latent vectors via a Perceiver bottleneck, and injects context-dependent bias vectors directly into the intermediate encoder representations of the ASR model. Because both the language model and the backbone ASR model remain entirely frozen during training, BaLEEN operates as a plug-and-play adapter that incurs zero computational overhead at inference time when context biases are precomputed. We evaluate our method on a CTC-based ASR model using a Wikipedia-derived corpus with annotated named entities and synthetic speech. Experimental results demonstrate that BaLEEN reduces keyword miss rate by 8.7% on the test set relative to the unbiased baseline while simultaneously improving overall word error rate by 21% and character error rate by 28%.
发表机构
- Sakana AI
机构由 AI 辅助整理,请以论文原文为准。