步长k子采样:无需训练的Whisper音频令牌缩减方法
Stride-k Subsampling: Train-Free Audio Token Reduction for Whisper
查看机构详情
- Chung-Ang University(中央大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究提出无需训练的步长k子采样方法,可减少Whisper音频令牌,降低计算量与延迟,在多数ASR基准上仅产生小的WER损失,适用于Whisper及基于其的语音语言模型。
中文摘要 AI 辅助
Whisper通过固定的1500令牌编码器接口输出语音,目前已成为ASR解码器和基于Whisper的语音语言模型(SpeechLMs)的默认表示形式,但其冗余性在很大程度上未被研究。我们提出步长k子采样,这是一种确定性索引操作,会保留卷积主干或编码器Transformer之后的每第k个令牌。在五种Whisper规模上,k=2在两个位置均保持了基线词错误率(WER),中心核对齐(CKA)将这种稳定性归因于主干处的声学重叠以及编码器输出处注意力诱导的重新分布。在两个位置应用步长2可减少75%的音频令牌,总GFLOPs降低52%-58%,在大多数ASR基准上的WER成本较小,在更难的基准上成本更大。该相同配置可扩展到三个基于Whisper的SpeechLMs,在较强的基线上产生适度的准确率下降,在较弱的基线上产生较大的下降,同时端到端延迟降低19.6%-27.4%。步长k子采样无需训练或辅助计算,利用了Whisper预处理的冗余性,表明其音频令牌接口的容量超过了下游任务的需求。
英文摘要
Whisper exposes speech through a fixed 1500-token encoder interface, now a default representation for ASR decoders and Whisper-based speech language models (SpeechLMs), yet its redundancy remains largely unexamined. We propose stride-k subsampling, a deterministic indexing operation that retains every k-th token after the convolutional stem or encoder transformer. Across five Whisper scales, k=2 preserves baseline WER at both positions, with CKA attributing this stability to acoustic overlap at the stem and attention-induced redistribution at the encoder output. Applying stride-2 at both positions cuts audio tokens by 75% and total GFLOPs by 52-58%, with small WER costs on most ASR benchmarks and larger costs on harder ones. The same configuration extends to three Whisper-based SpeechLMs, yielding modest accuracy drops on stronger baselines and larger drops on weaker ones, while reducing end-to-end latency by 19.6-27.4%. Requiring no training or auxiliary computation, stride-k subsampling exploits Whisper's preprocessing redundancy, indicating that its audio-token interface carries more capacity than downstream tasks require.