发表机构
Hitachi, Ltd., Research & Development Group; Hitachi Advanced Systems Corporation(株式会社日立制作所研发集团; 日立先进系统株式会社)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出无训练上下文ASR框架,利用SpeechLLM定位错误跨度并选择性检索术语,减少词典查询并提升多领域识别性能。
AI 中文摘要
领域特定和低频术语的识别对自动语音识别(ASR)仍具挑战性。尽管上下文偏置能改善其识别,但直接提供大型术语词典会引入许多无关的偏置项。基于检索的上下文偏置通过从外部词典中选择候选术语来解决此问题,但查询许多已识别单词需要大量词典查找,且可能产生针对性差的候选词。我们提出一种无训练的上下文ASR框架,其中预训练的语音大语言模型(SpeechLLM)联合生成ASR假设并定位可能涉及领域特定术语的错误跨度。仅使用定位的跨度从外部术语词典中检索音系相似的术语。同一SpeechLLM随后基于第一遍假设和检索到的术语重新识别音频,无需任务特定模型训练。为评估跨领域的适用性,我们在医疗、空中交通管制和金融语音上评估该框架。所提方法大幅减少词典查询,同时提高相关术语候选的召回率和排名,并在所有三个领域中提升第二遍ASR性能。
英文摘要
Recognition of domain-specific and low-frequency terms remains challenging for automatic speech recognition (ASR). Although contextual biasing can improve their recognition, directly providing a large terminology dictionary introduces many irrelevant biasing terms. Retrieval-based contextual biasing addresses this issue by selecting candidate terms from an external dictionary, but querying many recognized words requires numerous dictionary lookups and may yield poorly targeted candidates. We propose a training-free contextual ASR framework in which a pretrained speech large language model (SpeechLLM) jointly generates an ASR hypothesis and localizes error spans likely to involve domain-specific terms. Only the localized spans are used to retrieve phonologically similar terms from an external terminology dictionary. The same SpeechLLM then re-recognizes the audio conditioned on the first-pass hypothesis and the retrieved terms, without task-specific model training. To assess applicability across domains, we evaluate the framework on medical, air traffic control, and financial speech. The proposed method substantially reduces dictionary queries while improving the recall and ranking of relevant terminology candidates and second-pass ASR performance across all three domains.