通过组自适应融合网络提升说话人验证的公平性
Improving fairness in speaker verification via Group-adapted Fusion Network
- The Pennsylvania State University(宾夕法尼亚州立大学)
- Amazon Alexa AI(亚马逊Alexa AI)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对说话人验证模型在性别不平衡数据下对少数群体不公平的问题,提出组自适应融合网络(GFN),通过组嵌入自适应与分数融合,实现总体及少数群体等错误率显著降低并减少群体间性能差异。
中文摘要 AI 辅助
现代说话人验证模型使用深度神经网络将话语音频编码为具有判别性的嵌入向量。在训练过程中,这些网络通常被优化以区分任意说话人。这种学习过程会使对精细语音特征的学习偏向于占主导地位的人口统计群体,从而导致不同群体之间出现不公平的性能差异。这一点在具有相似语音特征的代表性不足的人口统计群体中尤为明显。在本工作中,我们在具有不平衡性别分布的控制数据集上研究了说话人验证模型的公平性,提供了模型性能在代表性不足群体中受损的直接证据。为缓解这一差异,我们提出了组自适应融合网络(GFN)架构,这是一种基于组嵌入自适应和分数融合的模块化架构。我们表明,我们的方法通过整体及针对各群体提升说话人验证性能来缓解模型不公平性。在训练中群体表示不平衡的情况下,与基线相比,我们提出的方法实现了总体等错误率(EER)相对降低9.6%至29.0%,少数群体EER降低13.7%至18.6%,并且EER差异减少20.0%至25.4%。该方法也适用于说话人识别系统中其他类型的训练数据偏斜。
英文摘要
Modern speaker verification models use deep neural networks to encode utterance audio into discriminative embedding vectors. During the training process, these networks are typically optimized to differentiate arbitrary speakers. This learning process biases the learning of fine voice characteristics towards dominant demographic groups, which can lead to an unfair performance disparity across different groups. This is observed especially with underrepresented demographic groups sharing similar voice characteristics. In this work, we investigate the fairness of speaker verification models on controlled datasets with imbalanced gender distributions, providing direct evidence that model performance suffers for underrepresented groups. To mitigate this disparity we propose the group-adapted fusion network (GFN) architecture, a modular architecture based on group embedding adaptation and score fusion. We show that our method alleviates model unfairness by improving speaker verification both overall and for individual groups. Given imbalanced group representation in training, our proposed method achieves overall equal error rate (EER) reduction of 9.6% to 29.0% relative, reduces minority group EER by 13.7% to 18.6%, and results in 20.0% to 25.4% less EER disparity, compared to baselines. The approach is applicable to other types of training data skew in speaker recognition systems.