面向说话人验证的统一不确定性感知后端:评分、归一化与校准
A Unified Uncertainty-Aware Back-End for Speaker Verification: Scoring, Normalization, and Calibration
浏览论文内容
中文总结 AI 辅助
该研究针对说话人验证后端未传递不确定性的问题,提出含UAS-Norm、UQMF的统一不确定性感知后端,在ECAPA-TDNN与ResNet上降低了等错误率并提升了区分度。
中文摘要 AI 辅助
说话人验证后端通常结合相似度评分、分数归一化与校准。但真实话语提取的说话人嵌入因时长、噪声、信道变化等因素存在与试次相关的可靠性问题。现有不确定性感知方法主要改进说话人编码器或初始相似度分数,而估计的不确定性通常未传递到后续的归一化与校准环节。我们将每个话语表示为解释为后验均值的说话人嵌入,以及作为不确定性估计的协方差。我们提出统一的不确定性感知后端,包含不确定性感知余弦评分、不确定性感知AS-Norm(UAS-Norm)、不确定性感知质量度量函数校准(UQMF),整个流程中纳入协方差信息以调整分数缩放、 cohort统计、归一化分数组合及校准特征。基于ECAPA-TDNN与ResNet的实验显示,两种架构均实现了一致的等错误率(EER)降低及目标-非目标区分度提升。
英文摘要
Speaker verification back-ends commonly combine similarity scoring, score normalization, and calibration. However, speaker embeddings extracted from real-world utterances have trial-dependent reliability because of factors such as duration, noise, and channel variation. Existing uncertainty-aware methods primarily improve the speaker encoder or the initial similarity score, while the estimated uncertainty is typically not propagated through subsequent normalization and calibration. We represent each utterance by a speaker embedding, interpreted as a posterior mean, together with its covariance as an uncertainty estimate. We present a unified uncertainty-aware back-end comprising uncertainty-aware cosine scoring, uncertainty-aware AS-Norm (UAS-Norm), and uncertainty-aware Quality Measure Function calibration (UQMF). Covariance information is incorporated throughout this pipeline to adjust score scaling, cohort statistics, normalized-score combination, and calibration features. Experiments with ECAPA-TDNN and ResNet show consistent EER reductions and improved target--non-target separation across both architectures.
发表机构
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。