为效率剪枝,以公平为代价:剪枝后的语音-大语言模型中的群体差异
Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs
- University of Essex(埃塞克斯大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究系统评估剪枝对语音-LLM在不同人口群体上的影响,发现剪枝加剧群体间WER差距,并建议部署时以最差群体错误率为明确标准。
AI中文摘要:
语音-大语言模型(Speech-LLMs)运行成本高昂,因此压缩对于实际部署至关重要。然而,压缩模型通常根据总体词错误率(WER)进行选择,这可能掩盖剪枝对不同人口群体(demographic groups)的影响。在本工作中,我们系统地研究了音频编码器剪枝对SLAM-ASR在不同人口群体上的效果。使用Fair-Speech和Common Voice数据集,我们发现剪枝并非对所有人口群体影响相同;最佳与最差表现群体之间的差距成倍增加。这些差异出现在所有三种编码器规模中,但只有最大的模型最初将其隐藏在总体WER之后。LoRA适配改善了每个群体的WER,但对原本表现较好的群体改善更强,并在某些群体间扩大了差距。在Common Voice英语、丹麦语和荷兰语中,口音差距持续存在但未明显扩大,表明剪枝的公平性影响因数据集而异,必须直接测量。我们的研究结果表明,对于剪枝模型,部署决策应包含按群体划分的WER,并将表现最差群体的错误率作为明确标准。
英文摘要:
Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we systematically study the effect of audio encoder pruning on SLAM-ASR for different demographic groups. Using the Fair-Speech and Common Voice datasets, we found that the pruning does not affect all demographic groups equally; the gap between best- and worst-performing groups increases in fold. These disparities appear across all three encoder scales, but only the largest model initially hides them behind aggregate WER. LoRA adaptation improves WER for every group, but benefits groups already performing well more strongly and widens for certain groups. On Common Voice English, Danish, and Dutch, accent gaps persist but do not clearly widen, showing that the fairness effects of pruning vary across datasets and must be measured directly. Our findings suggest that for pruned models, deployment decisions should include per-group WER, with the worst-performing group's error rate as an explicit criterion.