arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Whisper-Flash:基于声学条件的并行草稿生成以加速 Whisper 解码

Whisper-Flash: Acoustically Conditioned Parallel Drafting for Faster Whisper Decoding

Huapeng Zhou, Huayu Wang, Junkai Wu, Kangqi Wang, Xinyu Wang

arXiv 2609.32869首次发表:更新:

AI 中文总结

Whisper-Flash 利用语音识别中“未写词已说出”的特性,通过两层草稿模型并行生成八个词元,在 LibriSpeech 上实现 3.16 倍加速且输出不变,消融显示直接访问音频是关键。

AI 中文摘要

Whisper 是一种广泛使用的用于语音识别的编码器-解码器模型。其编码器在一次并行传递中读取话语,但其解码器一次一个词元地生成转录文本,这主导了推理时间。推测解码在不改变其输出的情况下缩短了此类循环:一个小的草稿模型猜测几个即将到来的词元,原始模型在一次前向传递中验证它们全部。我们提出了 Whisper-Flash,一个基于语音识别特性的两层草稿模型:尚未写出的单词已经被说出。它读取 Whisper 的编码音频和已接受的解码器状态,并在一次前向传递中提出八个词元。在完整的 LibriSpeech 测试集上,Whisper-Flash 每秒处理的音频量是贪心解码的 $3.16\ imes/2.85\ imes$,且输出相同,并且在批处理大小高达 96 以及温度采样下仍然更快。消融实验表明,直接访问音频最为重要。

英文摘要

Whisper is a widely used encoder-decoder model for speech recognition. Its encoder reads an utterance in one parallel pass, but its decoder writes the transcript one token at a time, which dominates inference time. Speculative decoding shortens such loops without changing their output: a small drafter guesses several upcoming tokens, and the original model verifies them all in one forward pass. We present Whisper-Flash, a two-layer drafter built on a property of speech recognition: the words still to be written have already been spoken. It reads Whisper's encoded audio and accepted decoder states and proposes eight tokens in a single forward pass. On the complete LibriSpeech test sets, Whisper-Flash processes $3.16\times/2.85\times$ as much audio per second as greedy decoding with identical outputs, and it remains faster at batch sizes up to 96 and under temperature sampling. Ablations show that direct access to the audio matters most.

Comments8 pages, 2 figures, 11 tables, including an appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑