发表机构
People Make Things(People Make Things)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出TRIAD审计网格和ORCA修复适配器,针对音频理解模型中的语音语义泄漏导致的人口统计公平性差异,通过对比头和正交惩罚将泄漏降低72%并大幅缩小差距。
AI 中文摘要
语音技术会惩罚某些声音:对于黑人说话者,识别错误率几乎是两倍,而对于第二语言口音和老年说话者,准确率也会下降。我们引入了TRIAD,一个审计网格,通过可控文本转语音跨越120个文本、24个渲染的人口统计语音档案(性别、年龄组、口音)和十种表达风格,将感知的人口统计属性与内容和情感隔离开来。对于十个开放权重编码器,我们定义了轴保真度泛函、轴子空间之间的主角度泄漏以及组条件差距;一个命题证明了平均探针差异随着我们测量的相同聚合语音语义泄漏$\Lambda$而增长,一个推论表明峰值泄漏会在活跃区域内强制产生最坏情况差异。测量的均方探针差异跟踪$\Lambda$(Pearson r = 0.93),并且一个黑盒协议在两种闭源模型中暴露了相同的特征。ORCA,一个结合了轴特定对比头、正交性惩罚和组平衡采样的适配器,将泄漏减少了72%,并将差距大致减半。
英文摘要
Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit grid crossing 120 texts, 24 rendered demographic voice profiles (gender, age band, accent), and ten expressive styles via controllable text-to-speech, isolating perceived demographic attributes from content and affect. For ten open-weights encoders we define axis-fidelity functionals, principal-angle leakage between axis subspaces, and group-conditional gaps; a proposition proves that average probe disparity grows with the same aggregate voice-semantic leakage $Λ$ we measure, and a corollary shows that peak leakage forces worst-case disparity inside an active region. The measured mean-square probe disparity tracks $Λ$ (Pearson r = 0.93), and a black-box protocol exposes the same signature in two closed-source models. ORCA, an adapter combining axis-specific contrastive heads, an orthogonality penalty, and group-balanced sampling, cuts leakage 72% and roughly halves the gaps.
CommentsAccepted by IEEE SLT 2026