AI 中文总结
提出从非流式ASR-LLM提取对齐路径并蒸馏给流式模型,改善对齐一致性,相对错误率降低16.6%。
AI 中文摘要
本文提出了一种面向流式自动语音识别(ASR)与大语言模型(LLM)的对齐路径蒸馏框架。交错式流式ASR-LLM使用来自对齐模型(如通过连接主义时间分类(CTC)训练的模型)的强制对齐(FA)来构建语音-文本训练序列。然而,从独立声学模型获得的对齐可能与基于LLM的ASR所学到的对齐不一致。这促使我们将非流式ASR-LLM的对齐信息迁移到流式识别中。具体而言,我们从非流式教师模型的软文本-音频注意力中提取单调对齐路径,并利用它们构建交错训练序列。该框架还包括logit和隐藏状态蒸馏,以学习教师模型的输出分布和内部表示。实验结果表明,在不使用logit或隐藏状态蒸馏的情况下,使用教师派生的对齐路径进行训练相比使用强制对齐训练,相对错误率降低了5.2%。当两个模型都使用logit和隐藏状态蒸馏时,教师派生的对齐使相对错误率降低了3.9%,平均发射延迟相似,但闪烁更高。完整框架相比使用强制对齐且无logit或隐藏状态蒸馏的训练,实现了16.6%的相对错误率降低。
英文摘要
In this paper, we propose an alignment-path distillation framework for streaming automatic speech recognition (ASR) with large language models (LLMs). Interleaved streaming ASR-LLMs use forced alignments (FA) from alignment models, such as those trained with connectionist temporal classification (CTC), to construct speech-text training sequences. However, alignments obtained from a separate acoustic model may be inconsistent with those learned by LLM-based ASR. This motivates us to transfer alignment information from a non-streaming ASR-LLM to improve streaming recognition. Specifically, we extract monotonic alignment paths from a non-streaming teacher's soft text-audio attention and use them to construct interleaved training sequences. The framework also includes logit and hidden-state distillation to learn from the teacher's output distributions and internal representations. Experimental results show that, without logit or hidden-state distillation, training with teacher-derived alignment paths achieves a 5.2% relative error rate reduction compared with training using forced alignments. When both models use logit and hidden-state distillation, teacher-derived alignments yield a 3.9% relative error rate reduction, with similar mean emission latency but higher flicker. The complete framework achieves a 16.6% relative error rate reduction compared with training using forced alignments without logit or hidden-state distillation.
Comments5 pages, 1 figure, 2 tables