AI 中文总结
该研究对比了基于Whisper的上下文偏置方法与语音LLMs在识别ASR中生词的性能,发现前者可大幅降低偏置词错误率,后者在朗读语音上表现好但泛化性差,最终为方法选择提供了权衡依据。
AI 中文摘要
识别训练数据中稀缺的生词和稀有词——包括命名实体、首字母缩略词、领域特定特殊词汇及其他项目——仍是自动语音识别(ASR)的关键挑战。我们对比两种应对策略:上下文偏置方法,即扩展ASR模型以便在推理过程中可提供单词列表;以及直接用上下文提示的语音大语言模型(LLMs)。我们基于Whisper的两种上下文偏置方法,针对三种语音LLMs,在朗读语音与非朗读语音上进行评估,报告了偏置、无偏置及整体词错误率(WER)。上下文偏置方法可将偏置WER相对降低最多88%,同时对其他单词几乎无影响;语音LLMs在朗读语音上表现出色,但对非朗读语音的泛化能力较差,且对干扰项数量和提示词顺序敏感。我们对由此产生的权衡进行了表征,以指导方法选择。
英文摘要
Recognizing new and rare words - named entities, acronyms, domain specific special words, and other items scarce in training data - remains a key challenge for automatic speech recognition (ASR). We compare two strategies for this: context biasing methods, where an ASR model is extended such that during inference a word list can be supplied, and speech large language models (LLMs) prompted with context directly. We evaluate two context biasing methods based on Whisper against three speech LLMs across read and non-read speech, reporting biased, unbiased, and overall word error rate (WER). The context biasing methods cut biased WER by up to 88% relative while leaving other words largely unaffected. Speech LLMs excel on read speech but generalize less well to non-read speech, and prove sensitive to distractor count and prompt word order. We characterize the resulting trade-offs to guide method selection.