发表机构
Maastricht University(马斯特里赫特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过探针实验发现,语音模型中的线性可读说话人属性方向虽能改变表征,但无法有效减少群体WER差距,表明表征引导并非可靠的公平性修复方法。
AI 中文摘要
自动语音识别(ASR)系统在不同说话人群体间表现出不等的错误率,这促使研究者对其内部表征进行干预。我们探究了从预训练ASR编码器中可线性读取的、与说话人相关的属性是否能提供有用的方向,以减少群体词错误率(WER)差距。在Common Voice和语音口音档案上,针对Whisper-medium、HuBERT-large和Wav2Vec2-large,我们对每个编码器层探测了基于元数据的性别、年龄和母语/口音标签;构建了质心和探针派生方向;在选定层注入这些方向;并将下游探针轨迹与匹配的WER变化进行比较。性别标签高度可解码(最佳宏F1为0.924--0.941),母语/口音标签也高于随机水平(0.544--0.696),而年龄较弱(0.354--0.397)。在22个后选择的重复运行中,有9个的95%配对自助区间完全低于零,但每个绝对源群体WER降低均低于0.7个百分点。相反,局部目标类别探针率可以从8.09%上升到99.87%,而WER却恶化。因此,线性可读性既不是因果使用的证据,也不是可靠的缓解方法。我们的结果促使在表征、传播和任务层面联合评估语音偏见干预措施。
英文摘要
Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representations. We ask whether speaker-linked attributes that are linearly readable from pretrained ASR encoders yield useful directions for reducing group word-error-rate (WER) gaps. Across Whisper-medium, HuBERT-large, and Wav2Vec2-large on Common Voice and the Speech Accent Archive, we probe every encoder layer for metadata-derived sex/gender, age, and native/accent labels; construct centroid and probe-derived directions; inject them at selected layers; and compare downstream probe trajectories with matched WER changes. Sex labels are highly decodable (best macro-F1 0.924--0.941), native/accent labels are also above chance (0.544--0.696), and age is weaker (0.354--0.397). Of 22 post-selected reruns, nine have 95% paired-bootstrap intervals entirely below zero, yet every absolute source-group WER reduction is below 0.7 percentage points. Conversely, a local target-class probe rate can rise from 8.09% to 99.87% while WER worsens. Linear readability is therefore neither evidence of causal use nor a reliable mitigation method. Our results motivate evaluating speech-bias interventions jointly at representation, propagation, and task levels.
CommentsAccepted at IMPACT-SPEECH 2026