发表机构
The Ohio State University; Air Force Research Laboratory(俄亥俄州立大学; 空军研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示 Whisper 模型训练后压缩(剪枝、量化)会加剧人口统计群体间的时间税差异,扩大词错误率差距,而蒸馏可缩小差距,表明全精度单次审计不足以评估部署时公平性。
AI 中文摘要
自动语音识别模型在全精度下接受人口统计公平性审计,但实际投入生产的模型已经过量化、剪枝和蒸馏。我们探究训练后权重压缩(其改变的是模型权重而非音频信号或其特征表示)是否会在不同人口统计群体间重新分配错误负担。在 Whisper 系列模型上,针对 Fair-Speech、Common Voice 25 和 AfriSpeech-200 数据集,对 Whisper-large-v3 进行 50% 的 Wanda 剪枝显著扩大了 Fair-Speech 上黑人/非裔美国人群体与亚裔群体之间的时间税差异:服务最差与最佳群体之间的绝对词错误率差距增加了一倍以上;假设每次转录错误需要五秒的纠正努力,则每分钟语音的纠正时间从 30 秒增加到 64 秒。这一 +111% 的相对增长对假设的每次错误成本不敏感,在音频质量控制下依然存在,且仅被波束搜索解码部分缓解,后者仍留下 +86% 的增长。在边缘模型规模下,INT4 HQQ 量化将西非口音上的灾难性转录循环放大了五到七倍。相比之下,蒸馏在 27 个评估设置(教师-学生配对、精度和数据集)中的 21 个中缩小了人口统计差距,例外情况集中在单一模型对上。我们将 Choi 和 Choi (2025) 的时间税概念转化为定量指标,并表明对全精度模型的单次快照公平性审计无法捕捉压缩给本已边缘化的说话者带来的部署时负担。
英文摘要
Automatic speech recognition models are audited for demographic fairness at full precision, yet the models that ship to production have been quantized, pruned, and distilled. We ask whether post-training weight compression, which alters model weights rather than the audio signal or its feature representation, redistributes error burden across demographic groups. Across the Whisper family on Fair-Speech, Common Voice 25, and AfriSpeech-200, 50% Wanda pruning of Whisper-large-v3 sharply widens the Black/AA-vs-Asian temporal-taxation differential on Fair-Speech: the absolute word-error-rate gap between the worst- and best-served groups more than doubles; at an assumed cost of five seconds of correction effort per transcription error this is a rise from 30 to 64 seconds of correction time per minute of speech. This +111% relative increase is invariant to the assumed per-error cost, survives an audio-quality control, and is only partly mitigated by beam-search decoding, which still leaves an +86% increase. At edge model size, INT4 HQQ quantization compounds catastrophic transcript loops on West African accents by factors of five to seven. Distillation, by contrast, narrows demographic gaps in 21 of 27 evaluated settings (teacher-student pair, precision, and dataset), with the exceptions concentrated on a single model pair. We cast the temporal-taxation construct of Choi and Choi (2025) as a quantitative metric, and show that single-snapshot fairness audits on full-precision models do not capture the deployment-time burden that compression places on already-marginalized speakers.
CommentsAccepted to IMPACT-SPEECH @ EMNLP 2026