arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13013cs.AIcs.SD

使用冻结离散扩散语言模型的音频原生语音识别

Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model

Harsha Vardhan Khurdula, Abhinav Kumar Singh, Yoeven D Khemlani, Vineet Agarwal

首次发表
浏览论文内容

中文总结 AI 辅助

探讨离散扩散语言模型能否用于语音识别,训练音频原生接口,用冻结Whisper编码器等,约42M参数。自然训练目标遇问题,连接主义时间分类损失打破僵局,模型在LibriSpeech test-clean上字错误率6.6%,能并行转录多种语言。

中文摘要 AI 辅助

自动语音识别由一次发出一个令牌的自回归解码器主导。我们探讨离散扩散语言模型能否取而代之,通过少量去噪步骤并行优化整个转录本。我们为DiffusionGemma训练了一个音频原生接口,它是一个26B专家混合模型,通过均匀、随机令牌离散扩散生成文本。冻结的Whisper编码器提供声学特征,轻量级投影仪将其映射到模型嵌入空间,低秩适配器让冻结主干关注新模态。训练了约42M参数,占主干的0.16%。自然训练目标无法将音频接地,通过冻结输出头应用的连接主义时间分类损失打破了僵局。结果模型在LibriSpeech test-clean上的字错误率达到6.6%,无论话语长度如何,大约在八个并行步骤中进行转录,并使用在六种语言上训练的单个适配器,我们在此对英语、印地语和普通话进行了评估。

英文摘要

Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps. We train an audio-native interface for DiffusionGemma, a 26B mixture-of-experts model that generates text by uniform, random-token discrete diffusion rather than the absorbing-mask scheme common to recent diffusion language models. A frozen Whisper encoder supplies acoustic features, a lightweight projector maps them into the model embedding space, and low-rank adapters let the frozen backbone attend to the new modality. About 42M parameters are trained, which is 0.16 percent of the backbone. We find that the natural training objectives fail to ground the audio because their gradient reaches the projector only through attention that has already dismissed it. A connectionist temporal classification loss applied through the frozen output head breaks this deadlock. The resulting model reaches 6.6 percent word error rate on LibriSpeech test-clean, transcribes in roughly eight parallel steps regardless of utterance length, and uses a single adapter trained on six languages, which we evaluate here on English, Hindi, and Mandarin.

发表机构

  • Interfaze AI(Interfaze人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑