arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27980cs.CLcs.LG

减少六层:基于无标签恢复的Whisper编码器剪枝

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

Rasmus Aagaard, Nicki Skafte Detlefsen

首次发表
浏览论文内容

中文总结 AI 辅助

针对Whisper编码器剪枝未广泛采用的问题,提出按逐层剔除的WER变化排序移除六层(占18.5%),并用无标签数据蒸馏恢复性能,平均WER从零样本21.9%降至20.1%。

中文摘要 AI 辅助

对大型预训练基于Transformer的自动语音识别(ASR)模型(如OpenAI的Whisper)进行剪枝已得到广泛采用,因为剪枝解码器带来了显著的端到端转录加速。例如,{\ t whisper-large-v3-turbo}变体将解码器从32层减少到4层,而Distill-Whisper同样将解码器减少到仅2层。尽管已有一些研究关注减小编码器规模,但尚无方法得到广泛采用。这可能是因为需要自定义推理实现才能利用压缩模型。我们提出了一种方法,通过逐层剔除(leave-one-layer-out)对词错误率(WER)的变化来对编码器层进行排序。移除导致变化最小的六层,对应编码器堆栈的$18.5\%$。剪枝后的模型无需自定义推理代码,因为它只是层数更少的更浅编码器。我们进一步使用无标签的单语语音数据进行蒸馏,以恢复由零样本层剪枝引起的性能下降。蒸馏后,四种语言的平均WER从零样本的$21.9\%$降至$20.1\%$,而基线为$18.2\%$。我们发布了所有代码(https://github.com/rasgaard/whisper-encoder-layer-prune)和剪枝模型(https://huggingface.co/rasgaard/whisper-large-v3-turbo-encoder-pruned)。

英文摘要

Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the {\tt whisper-large-v3-turbo} variant reduced the decoder from 32 to 4 layers, while Distill-Whisper similarly reduced the decoder to only 2 layers. Although some attention has been put towards reducing the size of the encoder, no approach has seen wide adoption. This could be due to the need for custom inference implementations to take advantage of the compressed model. We present an approach that ranks encoder layers by the leave-one-layer-out change in Word Error Rate (WER). The six layers that cause the least change are removed, corresponding to $18.5\%$ of the encoder stack. The pruned model requires no custom inference code as it is simply a more shallow encoder with fewer layers. We further distill using unlabeled monolingual speech data to recover performance degradation caused by the zero-shot layer pruning. Mean WER across four languages increases to $20.1\%$ after distillation, compared to $21.9\%$ zero-shot, going from a baseline of $18.2\%$. We release all of our code (https://github.com/rasgaard/whisper-encoder-layer-prune) and the pruned model (https://huggingface.co/rasgaard/whisper-large-v3-turbo-encoder-pruned).

发表机构

  • Technical University of Denmark(丹麦技术大学)
  • Laerdal Medical(挪度医疗)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑