熵感知逻辑回归用于大规模说话人识别系统的融合
Entropy-aware logistic regression for fusion of large-scale speaker recognition systems
查看机构详情
- LIA, Avignon University(阿维尼翁大学LIA实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对传统逻辑回归融合忽略话语可靠性变化的问题,提出熵感知逻辑回归,结合系统互补性与话语不确定性,提升大规模说话人识别性能。
中文摘要 AI 辅助
基于逻辑回归的分数级融合在说话人识别中被广泛用于组合互补系统。然而,传统方法分配固定的系统相关系数,并未明确考虑单个注册和测试话语可靠性的变化。借鉴近期关于基于深度学习的说话人识别模型熵的研究,本研究将不确定性成分纳入融合过程。通过利用系统级互补性和话语相关不确定性,该方法在涉及语音信号高度可变特征的大规模说话人识别任务中实现了稳健性能。这些结果表明,模型熵信息在大规模场景中提供了有价值的互补线索。
英文摘要
Score-level fusion based on logistic regression is widely used in speaker recognition to combine complementary systems. However, conventional approaches assign fixed system-dependent coefficients and do not explicitly account for variations in the reliability of individual enrollment and test utterances. Drawing on recent research on the entropy of deep learning-based speaker recognition models, this study incorporates an uncertainty component into the fusion process. By exploiting both system-level complementarity and utterance-dependent uncertainty, the method achieves robust performance in large-scale speaker recognition tasks that involve highly variable characteristics of the speech signal. These results demonstrate that model-entropy information provides a valuable complementary cue in large-scale scenarios.