语音块影响:用于剪枝语音大语言模型的组件特定层评分
Speech Block Influence: Component-Specific Layer Scoring for Pruning Speech LLMs
浏览论文内容
中文总结 AI 辅助
提出语音块影响(SBI)框架,通过编码器与解码器组件特定评分,提升语音大语言模型层剪枝的鲁棒性与效率。
中文摘要 AI 辅助
语音大语言模型在资源受限环境中的部署成本高昂。层剪枝可以降低这一成本,但现有的评分指标难以迁移到语音大语言模型:它们假设解码器专用架构且具有同质词元序列,而语音大语言模型增加了编码器和适配器组件,并处理多模态序列。我们提出语音块影响(SBI),这是首个专为语音大语言模型剪枝设计的层重要性评分框架,包含两个组件特定分数:SBI-Enc在适配器输出处衡量编码器层移除的影响,以更好地反映下游影响;SBI-Dec仅针对文本词元位置计算逐层输入-输出相似度,以避免音频词元主导。在三个语音大语言模型上,SBI提升了剪枝鲁棒性,在较高剪枝率下实现了更强的编码器性能,并通过评分文本词元而非音频主导的完整序列,实现了更可靠的解码器层选择。我们进一步发现,仅文本校准产生的解码器排名与语音-文本校准的排名高度相关,表明存在一种更便宜的替代方案来衡量解码器层重要性。
英文摘要
Speech LLMs are costly to deploy in resource-constrained settings. Layer pruning can cut this cost, but existing scoring metrics transfer poorly to speech LLMs: they assume a decoder-only architecture with homogeneous token sequences, whereas speech LLMs add encoder and adapter components and process multimodal sequences. We propose Speech Block Influence (SBI), the first layer-importance scoring framework designed for speech LLM pruning that consists of two component-specific scores: SBI-Enc measures the effect of encoder-layer removal at the adapter's output to better reflect downstream impact; SBI-Dec measures layer-wise input-output similarity over text-token positions only to avoid audio-token dominance. Across three speech LLMs, SBI improves pruning robustness, with stronger encoder performance at higher pruning rates and more reliable decoder layer selection by scoring text tokens rather than the audio-dominated full sequence. We further find that text-only calibration yields decoder rankings highly correlated with those from speech-text calibration, suggesting a cheaper alternative to measure decoder layer importance.
发表机构
- The University of Western Australia(西澳大利亚大学)
机构由 AI 辅助整理,请以论文原文为准。