发表机构
University of Luxembourg; Radio Télévisioun Lëtzebuerg (RTL)(卢森堡大学; 卢森堡广播电视公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对TidyVoice 2026挑战赛中的跨语言说话人验证问题,从官方基线出发,将干扰属性投影作为嵌入空间的简单语言归一化步骤,通过估计语言子空间并投影嵌入,降低了开发集等效错误率,提升了评估分数。
AI 中文摘要
跨语言不匹配仍然是现代说话人验证整体性能下降的关键原因。TidyVoice2026挑战赛针对此情况进行文本无关验证,有40种语言的3666名训练和808名开发说话人,38种未见语言的2200名评估说话人,测试时无语言标签。从在VoxBlink2和VoxCeleb2上预训练并在TidyVoice上微调的官方SimAM-ResNet34基线开始,我们将干扰属性投影(NAP)作为嵌入空间中的简单语言归一化步骤重新审视。我们从跨语言同说话人差异估计一个紧凑语言子空间,并在使用自适应对称分数归一化进行余弦评分之前将嵌入投影到其正交补空间上。这将开发集的等效错误率从余弦法的2.97%和AS-Norm法的2.70%降至2.18%,并产生了8.40的Codabench评估分数,表明简单的后端语言归一化可以与更复杂的系统相媲美。
英文摘要
Cross-lingual mismatch remains a key source of overall degradation in modern speaker verification. The TidyVoice2026 Challenge targets this setting with text-independent verification, comprising 3,666 training and 808 development speakers in 40 languages and 2,200 evaluation speakers in 38 unseen languages, without language labels at test time. Starting from the official SimAM-ResNet34 baseline pretrained on VoxBlink2 and VoxCeleb2 and fine-tuned on TidyVoice, we revisit Nuisance Attribute Projection (NAP) as a simple language-normalization step in the embedding space. We estimate a compact language subspace from cross-language same-speaker differences and project embeddings onto its orthogonal complement before cosine scoring with Adaptive Symmetric score normalization. This reduces development EER from 2.97\% with cosine and 2.70\% with AS-Norm to 2.18\% and yields a Codabench evaluation score of 8.40, showing that simple back-end language normalization can rival more complex systems.
Comments5 pages, 2 figures (submitted to the TidyVoice challenge colocated with Interspeech2026)