AI 中文总结
本文提出基于非自回归大语言模型的架构,为给定转写文本添加时间戳和说话人信息,较自回归模型更快更准,实现最先进的时间戳准确性和最佳cpWER。
AI 中文摘要
时间戳和说话人归属是语音识别的有益补充,能够生成丰富的文本转写。这些信息既可以在转录过程中提取,也可以对齐到给定的转写文本上。本文提出了一种基于非自回归大语言模型架构的模型,用于为给定的转写文本添加时间戳和说话人信息。与由相似组件构建的自回归模型相比,该模型更加准确,并且注释给定转写文本的速度快一到两个数量级。与其他模型相比,我们的模型在时间戳准确性上达到了最先进水平,并在说话人归属上取得了最佳的cpWER。
英文摘要
Timestamps and speaker attribution are useful additions to speech recognition, creating a rich text transcript. This information can either be extracted during transcription or aligned to a given transcript. In this paper we present models that add timestamps and speaker information to a given transcript using a non-autoregressive LLM-based architecture. Compared to an autoregressive model built from similar components, the models are more accurate and annotate a given transcript one to two orders of magnitude faster. Compared to other models, our models achieve state-of-the-art timestamp accuracy and the best cpWER for speaker attribution.
CommentsSubmitted to ICASSP 2027