发表机构
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Triage方法,在语言模型运行前通过线性预测音频令牌的注意力排序进行剪枝,在转录和多项选择任务上优于基线,并显著提升上下文窗口容量。
AI 中文摘要
大型音频语言模型(LALM)将一分钟的语音转换为750至1500个令牌,并对每个令牌进行预填充。图像令牌剪枝通常在语言模型的前几层之后进行,因为图像令牌在那里获得的注意力很少。音频令牌在该处获得的注意力要多得多,且其排序远未定型,因此音频需要在语言模型运行之前进行排序。令人惊讶的是,音频令牌在语言模型中将要获得的注意力,已经可以从其编码器输出中线性预测,且无需语言模型运行。一个以闭式形式拟合、无需标签的线性映射,在十三个LALM中的十一个上,以ρ≥0.69的相关性预测了这种全层注意力排序。我们的方法Triage,通过该预测剪除音频令牌,并在多项选择任务中,于第2层再次剪除,利用该处观察到的注意力修正预测。Triage在无需标签的情况下设定其压缩率,受两个预算约束,限制其输出与模型自身全音频输出的差异程度。在保守预算下,其词错误率和准确率与全音频相差在0.04以内。在激进预算下,Triage在所有十二个转录案例中均优于所有基线。在多项选择任务中,在2.2至5倍压缩下,其平均准确率比最强基线DART高出0.043。由于它在语言模型之前剪除,它将Qwen2.5-Omni-3B上下文窗口可容纳的音频从21.8分钟提高到约62分钟。在其最大压缩点,Triage让单个GPU能够为该模型服务4倍多的并发5分钟流。项目页面:此https URL
英文摘要
A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their ranking is still far from final, so audio needs a ranking before the language model runs. Surprisingly, the attention an audio token will receive across the language model is already linearly predictable from its encoder output, before the language model runs. A linear map, fitted in closed form without labels, predicts this all-layer attention ranking at $ρ\geq .69$ on eleven of thirteen LALMs. Our method, Triage, cuts audio tokens by this prediction and, on multiple choice, cuts again at layer 2, correcting the prediction with the attention observed there. Triage sets its compression without labels, under two budgets that limit how far its output may differ from the model's own full-audio output. At the conservative budget, its word error rate and accuracy stay within .04 of full audio. At the aggressive budget, Triage beats every baseline in all twelve transcription cases. On multiple choice, at 2.2-5x compression, it outperforms DART, the strongest baseline on average, by .043 in mean accuracy. Because it cuts before the language model, it raises the audio that fits in Qwen2.5-Omni-3B's context window from 21.8 to about 62 minutes. At its most compressive point, Triage lets one GPU serve 4x as many concurrent 5-minute streams of that model. Project page: https://audio-triage.github.io
Comments43 pages, 7 figures