arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于分解延迟流建模的LLM流式语音识别

Factorized Delayed Streams Modeling for LLM-based Streaming ASR

Tatsunari Takagi, Kai Washizaki, Atsushi Kojima, Lianbo Liu, Koki Nikaido, Yui Sudo

arXiv 2610.04333首次发表:更新:

发表机构

SB Intuitions; Toyohashi University of Technology(SB Intuitions; 丰桥技术科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出F-DSM,通过分解等待概率与词汇分布,移除ASR特定标记并跳过softmax,在提升流式ASR识别性能的同时降低内存和推理成本。

AI 中文摘要

延迟流建模(DSM)通过在共同时间线上对齐声学流和文本流,使基于大型语言模型(LLM)的流式自动语音识别(ASR)成为可能。DSM在LLM词汇表中添加填充标记<p>和词首标记<w>,并使用相同的softmax与正常文本标记一起预测它们。我们首先证明,在保持竞争性识别性能的同时,可以移除<w>。基于这一结果,我们提出分解延迟流建模(F-DSM),它将<p>的等待概率与原始LLM词汇表上的分布分离。这种分解将ASR特定标记从文本预测空间中移除,并允许在等待步骤跳过大规模词汇softmax。在日语自发语音语料库和LibriSpeech上的实验表明,F-DSM比DSM实现了更好的识别性能。它还大幅减少了GPU内存使用,同时保持相似的训练吞吐量,通过跳过softmax提供了小幅推理速度提升,并减少了DSM中观察到的纯文本困惑度下降。

英文摘要

Delayed Streams Modeling (DSM) enables LLM-based streaming automatic speech recognition (ASR) by aligning acoustic and text streams on a common timeline. DSM adds the padding token <p> and the word-start token <w> to the LLM vocabulary and predicts them together with normal text tokens using the same softmax. We first show that <w> can be removed while maintaining competitive recognition performance. Based on this result, we propose Factorized DSM (F-DSM), which separates the waiting probability for <p> from the distribution over the original LLM vocabulary. This factorization removes ASR-specific tokens from the text prediction space and allows the large-vocabulary softmax to be skipped on waiting steps. Experiments on the Corpus of Spontaneous Japanese and LibriSpeech show that F-DSM achieves better recognition performance than DSM. It also greatly reduces GPU memory use while maintaining similar training throughput, provides a small inference speed improvement through softmax skipping, and reduces the degradation in text-only perplexity observed with DSM.

CommentsSubmitted to IEEE ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑