arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38976cs.CL

超越单次运行的公平性:语音大语言模型适配中的训练种子变异性

Fairness Beyond a Single Run: Training-Seed Variability in Speech LLM Adaptation

Srishti Ginjala, Eric Fosler-Lussier, Srinivasan Parthasarathy

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示语音大语言模型适配中训练种子对公平性指标的影响超过压缩因子,建议报告多次运行结果以准确评估公平性。

中文摘要 AI 辅助

自动语音识别中的群体公平性差距几乎总是基于单次训练运行来报告的。我们在五个音频压缩因子和六个随机种子下对语音大语言模型的Q-former投影器和LoRA适配器进行微调,同时保持编码器、基础解码器、数据和解码固定不变,并在Common Voice和Fair-Speech上评估每次运行。在460小时的干净LibriSpeech数据上,种子对公平性指标的影响在大多数人口统计维度上超过了压缩的影响。一个平衡的3x3分解显示,Fair-Speech种族归一化差距变异的85.3%归因于种子,而压缩仅占8.3%(p = 0.009),尽管压缩在年龄和性别维度上解释了更多变异。留出的LibriSpeech词错误率在这些种子间波动0.04个百分点,而Common Voice波动8.57个百分点,因此这些并非失败运行,且该效应在控制准确率和dropout后依然存在。将适配集扩展到并多样化至960小时会减弱该效应,但并未消除它。在Fair-Speech种族维度上,两个单次运行系统之间的归一化差距必须超过0.30才能超出种子变异性。

英文摘要

Demographic fairness gaps in automatic speech recognition are almost always reported from a single training run. We fine-tune the Q-former projector and LoRA adapters of a speech LLM at five audio compression factors and six random seeds, holding the encoder, base decoder, data and decoding fixed, and evaluate every run on Common Voice and Fair-Speech. At 460 h of clean LibriSpeech, the seed moves fairness metrics more than compression does on most demographic axes. A balanced 3x3 decomposition attributes 85.3% of the variation in Fair-Speech ethnicity normalized gap to the seed against 8.3% to compression (p = 0.009), though compression explains more on age and gender. Held-out LibriSpeech word error rate spreads by 0.04 points across those seeds while Common Voice spreads by 8.57, so these are not failed runs, and the effect survives controlling for accuracy and dropout. Scaling and diversifying the adaptation set to 960 h damps the effect but does not remove it. On Fair-Speech ethnicity, two single-run systems must differ by more than 0.30 in normalized gap to exceed seed variability.

发表机构

  • The Ohio State University(俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑