arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11272eess.AS

利用注册声纹库作为证据:面向目标说话人标注的仿射不变分数校准

Using the enrollment gallery as evidence: affine-invariant score calibration for target speaker tagging

Hee-Soo Heo, Minjae Lee, Youngki Kwon, Bong-Jin Lee

AI总结:

该研究针对目标说话人标注问题,提出声纹库亲和度验证的无 cohort 仿射不变分数校准方法,在合成基准和会议语料库上提升了标注准确率,结合传统归一化后效果更佳。

AI中文摘要:

目标说话人标注(TST)是将已注册的说话人身份分配给多说话人录音的分时段语音片段。其他已注册说话人的注册语音是校准每个验证决策的理想证据来源:它们与测试材料处于同一部署域,且无需外部 cohort 集。然而,我们发现这些注册语音无法作为传统分数归一化的 cohort。注册语音往往按录音会话聚类,因此每个说话人的 cohort 统计量反映的是注册语音与声纹库其余部分的接近程度,而非冒名者行为,归一化会进而拒绝整个说话人。我们提出了声纹库亲和度验证,这是一种无 cohort 的校准方法,通过两个仿射不变项(分数离散度和注册语音轮廓偏差)利用声纹库,因此不受此类每个说话人偏移的影响。在合成基准和内部会议语料库上,该方法在使用和不使用传统分数归一化的情况下均提高了标注准确率,且与传统分数归一化结合时进一步提升了性能。

英文摘要:

Target speaker tagging (TST) assigns enrolled speaker identities to diarized segments of multi-speaker recordings. The enrollment utterances of the other registered speakers are an attractive source of evidence for calibrating each verification decision: they share the deployment domain of the test material and require no external cohort set. We show, however, that they cannot serve as the cohort of conventional score normalization. Enrollments tend to cluster by recording session, so the per-speaker cohort statistics reflect enrollment proximity to the rest of the gallery rather than impostor behavior, and normalization then rejects entire speakers. We propose gallery affinity verification, a cohort-free calibration that exploits the gallery through two affine-invariant terms, score dispersion and enrollment-profile deviance, and is therefore immune to such per-speaker shifts. On a synthetic benchmark and an in-house meeting corpus, it improves tagging accuracy with and without conventional score normalization and adds further gains when combined with it.

补充信息

↑