ParaASR:面向快速长上下文基于大语言模型的语音识别的多令牌预测
ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition
另 3 家 · 查看机构详情
- StepFun(阶跃星辰)
- NTU(南洋理工大学)
- PKU(北京大学)
- UNSW(新南威尔士大学)
- SJTU(上海交通大学)
- USTC(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
ParaASR是一款基于大语言模型的ASR系统,通过多令牌预测技术,在保持低错误率、低延迟的同时,支持32K长上下文及30分钟音频的快速转录,兼顾识别质量、速度与上下文长度。
中文摘要 AI 辅助
音频编码器-大语言模型解码器架构已成为现代自动语音识别(ASR)的主流范式,通过大规模语言建模提升转录质量。然而,自回归解码的成本随解码器规模增大而升高,在识别质量与服务延迟间形成了根本权衡。本文认为该权衡并非固有:与开放式文本生成不同,ASR输出紧密锚定输入语音信号,为高并行解码提供了天然归纳偏置。基于此,我们推出ParaASR,一款利用多令牌预测(MTP)的ASR系统,让4B规模的大语言模型解码器每前向步骤生成多个令牌。从公开的音频-语言基础模型出发,该模型先构建稳健的自回归识别器,再通过分阶段优化方案对齐5个未来令牌分支。推理时,它每步生成6个令牌的延续内容,仅将经验证的前缀纳入转录,保留了标准自回归解码的安全性。平均接受长度达生成6个令牌中的5.0个,证实语音的确定性结构使ASR成为多令牌解码的天然适用场景。ParaASR还保留原生32K上下文窗口,单次可转录长达30分钟的音频。在各类基准测试中,它在中文、英文及长文本评估上的平均错误率分别为2.97%、3.68%和3.70%,同时达到低至0.0053的实时因子(RTF)。这些结果表明,当未来令牌提议由声学信号锚定并经自回归验证保障时,解码器规模扩展、低延迟推理与长上下文转录无需成为相互竞争的目标。
英文摘要
Audio-encoder-LLM-decoder architectures have become the dominant paradigm for modern automatic speech recognition (ASR), improving transcription quality through large-scale language modeling. However, the cost of autoregressive decoding scales with decoder size, creating a fundamental trade-off between recognition quality and serving latency. We argue this trade-off is not inherent: unlike open-ended text generation, ASR outputs are strongly anchored to the input speech signal, providing a natural inductive bias toward high-parallelism decoding. Building on this, we introduce ParaASR, an ASR system that leverages Multi-Token Prediction (MTP) to let a 4B LLM decoder emit multiple tokens per forward step. Starting from a publicly available audio-language foundation, the model first establishes a robust autoregressive recognizer and then aligns five future-token branches through a staged optimization recipe. At inference, it proposes a six-token continuation per step and admits only the verified prefix into the transcript, preserving the safety of standard autoregressive decoding. The average accepted length reaches 5.0 out of 6 proposed tokens, confirming that the deterministic structure of speech makes ASR an especially natural setting for multi-token decoding. ParaASR further retains a native 32K-context window and transcribes up to 30 minutes of audio in a single pass. Across diverse benchmarks, it attains average error rates of 2.97%, 3.68%, and 3.70% on Chinese, English, and long-form evaluations, respectively, while reaching a real-time factor (RTF) as low as 0.0053. These results show that decoder scaling, low-latency inference, and long-context transcription need not be competing goals when future-token proposals are anchored by the acoustic signal and guarded by autoregressive verification.