AI 中文总结
针对LLM直接重评分ASR假设成本高的问题,本文提出无需训练的缓存LLM概率检索方法,在多数设置下优于1-pass ASR,是轻量有效的ASR适配方案。
AI 中文摘要
大语言模型(LLM)通过提供语言先验提升自动语音识别(ASR)性能,但直接对每个N-best假设进行重评分的成本很高。本文提出“缓存LLM概率检索”方法,该方法离线查询本地教师LLM,获取与ASR相关的上下文-目标对的下一个词元概率,在识别阶段通过缓存查找、回退策略及对显著未命中情况的可选评分来利用这些概率。该方法无需训练,可与现有识别器集成且无需修改声学模型。在39种设置中,缓存检索在28种设置下优于1-pass ASR,且非oracle错误更低。上下文长度分析显示,在上下文长度为8时收益达到峰值,表明缓存概率检索是一种有效且轻量的ASR适配方法,与生成错误修正(GER)或知识蒸馏(KD)所需的大量训练形成对比。
英文摘要
Large language models (LLMs) enhance automatic speech recognition (ASR) by providing linguistic priors; however, their direct rescoring is costly because it requires evaluating every N-best hypothesis. This paper introduces "cached LLM probability retrieval," which involves querying a local teacher LLM offline to obtain next-token probabilities for ASR-relevant context-target pairs. These probabilities are then utilized during recognition via cache lookups, backoff strategies, and optional scoring for significant misses. The method is training-free and can integrate with existing recognizers without requiring modifications to acoustic models. Evaluations across various ASR models reveal that cached retrieval outperforms 1-pass ASR in 28 of 39 settings and achieves lower non-oracle errors. Context length analysis indicates that benefits peak at a context length of 8, suggesting that cached probability retrieval is an effective and lightweight ASR adaptation method, in contrast to the heavy training required for Generative Error Correction (GER) or knowledge distillation (KD).
Commentsunder review