发表机构
lab260; BitmanagerAI; MTUCI(lab260; BitmanagerAI; MTUCI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ReDimNet2+,通过多语料库数据扩展、编解码器与波形增强、大间隔微调和图检索重排序,显著提升说话人验证的跨设备稳健性,大幅降低EER并改善检索性能。
AI 中文摘要
自动说话人验证必须在不同设备、房间和压缩流程中保持可靠性。我们提出了ReDimNet2+,它在七个公开语料库(63,934个说话人,约8,675小时)上扩展了紧凑型ReDimNet2骨干网络的训练。对VoxBlink2子集的分析揭示了预测频谱着色的偏移,这促使我们在多语料库训练中引入编解码器和波形增强,并结合大间隔微调(LMFT)和基于图的检索重排序。在所有模型使用随机4秒评估窗口的情况下,ReDimNet2+ LMFT将池化VoxCeleb1的EER从2.42%降低到0.82%,并将26种条件下的鲁棒性压力测试EER从7.21%降低到1.99%。在此共享本地协议下,它在VoxCeleb1-O上达到0.35%的EER,而最佳评估的WeSpeaker检查点为0.787%。在VoxBlink2检索子集上,重排序将最终模型的Pr@k从0.7413提升到0.7687。
英文摘要
Automatic speaker verification must remain reliable across devices, rooms, and compression pipelines. We present ReDimNet2+, which scales training of the compact ReDimNet2 backbone across seven public corpora (63,934 speakers, about 8,675 hours). Analysis of a VoxBlink2 subset reveals a shift in predicted spectral coloration, motivating codec and waveform augmentation alongside this multi-corpus training, large-margin fine-tuning (LMFT), and graph-based retrieval reranking. With random 4-second evaluation windows for all models, ReDimNet2+ LMFT reduces pooled VoxCeleb1 EER from 2.42% to 0.82% and a 26-condition robustness stress-test EER from 7.21% to 1.99%. Under this shared local protocol, it reaches 0.35% EER on VoxCeleb1-O versus 0.787% for the best evaluated WeSpeaker checkpoint. On a VoxBlink2 retrieval subset, reranking improves the final model's Pr@k from 0.7413 to 0.7687.
CommentsSubmitted to ICASSP 2027