arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29279cs.SD

ParaASR:面向快速长上下文基于大语言模型的语音识别的多令牌预测

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition

发表机构阶跃星辰 · 南洋理工大学 · 北京大学
另 3 家 · 查看机构详情
  • StepFun(阶跃星辰)
  • NTU(南洋理工大学)
  • PKU(北京大学)
  • UNSW(新南威尔士大学)
  • SJTU(上海交通大学)
  • USTC(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Qingjian Lin, Yuxin Li, Haoyang Zhang, Jun Chen, Yechang Huang, Feng Tian, Xie Li, Xiangyu Tony Zhang, Daijiao Liu, Yuxin Zhang, Jinglan Gong, Bo Zhao, Fei Tian… 展开作者

Qingjian Lin, Yuxin Li, Haoyang Zhang, Jun Chen, Yechang Huang, Feng Tian, Xie Li, Xiangyu Tony Zhang, Daijiao Liu, Yuxin Zhang, Jinglan Gong, Bo Zhao, Fei Tian, Xuerui Yang, Gang Yu, Xiangyu Zhang, Daxin Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

ParaASR是一款基于大语言模型的ASR系统,通过多令牌预测技术,在保持低错误率、低延迟的同时,支持32K长上下文及30分钟音频的快速转录,兼顾识别质量、速度与上下文长度。

中文摘要 AI 辅助

音频编码器-大语言模型解码器架构已成为现代自动语音识别(ASR)的主流范式,通过大规模语言建模提升转录质量。然而,自回归解码的成本随解码器规模增大而升高,在识别质量与服务延迟间形成了根本权衡。本文认为该权衡并非固有:与开放式文本生成不同,ASR输出紧密锚定输入语音信号,为高并行解码提供了天然归纳偏置。基于此,我们推出ParaASR,一款利用多令牌预测(MTP)的ASR系统,让4B规模的大语言模型解码器每前向步骤生成多个令牌。从公开的音频-语言基础模型出发,该模型先构建稳健的自回归识别器,再通过分阶段优化方案对齐5个未来令牌分支。推理时,它每步生成6个令牌的延续内容,仅将经验证的前缀纳入转录,保留了标准自回归解码的安全性。平均接受长度达生成6个令牌中的5.0个,证实语音的确定性结构使ASR成为多令牌解码的天然适用场景。ParaASR还保留原生32K上下文窗口,单次可转录长达30分钟的音频。在各类基准测试中,它在中文、英文及长文本评估上的平均错误率分别为2.97%、3.68%和3.70%,同时达到低至0.0053的实时因子(RTF)。这些结果表明,当未来令牌提议由声学信号锚定并经自回归验证保障时,解码器规模扩展、低延迟推理与长上下文转录无需成为相互竞争的目标。

英文摘要

Audio-encoder-LLM-decoder architectures have become the dominant paradigm for modern automatic speech recognition (ASR), improving transcription quality through large-scale language modeling. However, the cost of autoregressive decoding scales with decoder size, creating a fundamental trade-off between recognition quality and serving latency. We argue this trade-off is not inherent: unlike open-ended text generation, ASR outputs are strongly anchored to the input speech signal, providing a natural inductive bias toward high-parallelism decoding. Building on this, we introduce ParaASR, an ASR system that leverages Multi-Token Prediction (MTP) to let a 4B LLM decoder emit multiple tokens per forward step. Starting from a publicly available audio-language foundation, the model first establishes a robust autoregressive recognizer and then aligns five future-token branches through a staged optimization recipe. At inference, it proposes a six-token continuation per step and admits only the verified prefix into the transcript, preserving the safety of standard autoregressive decoding. The average accepted length reaches 5.0 out of 6 proposed tokens, confirming that the deterministic structure of speech makes ASR an especially natural setting for multi-token decoding. ParaASR further retains a native 32K-context window and transcribes up to 30 minutes of audio in a single pass. Across diverse benchmarks, it attains average error rates of 2.97%, 3.68%, and 3.70% on Chinese, English, and long-form evaluations, respectively, while reaching a real-time factor (RTF) as low as 0.0053. These results show that decoder scaling, low-latency inference, and long-context transcription need not be competing goals when future-token proposals are anchored by the acoustic signal and guarded by autoregressive verification.

补充信息

↑