基于小波提示调优和多模型集成的防伪语音验证
Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles
AI总结:
本文提出了一种结合小波提示调优和多模型集成的防伪语音验证方法,通过在VoxCeleb2和SpoofCeleb上训练,实现了低误检率和高泛化能力。
AI中文摘要:
本文描述了参加WildSpoof 2026挑战赛SASV部分的UZH-CL系统。该挑战旨在通过同时验证说话人身份和音频真实性来集成防御生成伪造攻击。我们提出了一种级联的防伪意识语音验证框架,该框架集成了小波提示调优的XLSR-AASIST反制措施和多模型集成。ASV组件使用了ResNet34、ResNet293和WavLM-ECAPA-TDNN架构,随后进行Z分数标准化和分数平均。在VoxCeleb2和SpoofCeleb上训练的系统获得了宏a-DCF为0.2017和SASV EER为2.08%。虽然系统在域内数据上实现了0.16%的伪造检测EER,但在未见数据集如ASVspoof5上的结果突显了跨域泛化的重要性。
英文摘要:
This paper describes the UZH-CL system submitted to the SASV section of the WildSpoof 2026 challenge. The challenge focuses on the integrated defense against generative spoofing attacks by requiring the simultaneous verification of speaker identity and audio authenticity. We proposed a cascaded Spoofing-Aware Speaker Verification framework that integrates a Wavelet Prompt-Tuned XLSR-AASIST countermeasure with a multi-model ensemble. The ASV component utilizes the ResNet34, ResNet293, and WavLM-ECAPA-TDNN architectures, with Z-score normalization followed by score averaging. Trained on VoxCeleb2 and SpoofCeleb, the system obtained a Macro a-DCF of 0.2017 and a SASV EER of 2.08%. While the system achieved a 0.16% EER in spoof detection on the in-domain data, results on unseen datasets, such as the ASVspoof5, highlight the critical challenge of cross-domain generalization.