arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

统一共享编码器在混合小波提示调优的防欺骗说话人验证中的应用

Unified Shared Encoder in Spoof-Aware Speaker Verification with Hybrid Wavelet Prompt Tuning

Aref Farhadipour, Srikanth Madikeri, Teodora Vukovic, Volker Dellwo, Petr Motlicek

arXiv 2610.08845首次发表:更新:

发表机构

University of Zurich; Idiap Research Institute; Brno University of Technology(苏黎世大学; Idiap研究所; 布尔诺理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出统一共享编码器结合混合小波提示调优,解决防欺骗说话人验证中效率与准确性的权衡,仅训练少量参数即超越双编码器级联,并在现实威胁下保持稳健。

AI 中文摘要

防欺骗说话人验证(SASV)必须确认说话人身份并确保语音是真实的,但当前系统往往顾此失彼。模块化级联系统准确但需运行两个大型自监督编码器,而端到端模型效率高但单个嵌入需服务于两个冲突目标时准确性下降。我们证明这种效率与专业化之间的权衡是可以避免的。单个冻结的W2V-BERT~2.0骨干网络通过深度混合小波提示调优(WPT)按任务适配,每个分支通过专用提示和任务头获得共享编码器的独立视图,仅在推理时合并分数。仅训练约700万参数,远少于强双编码器级联所需的参数,该系统在SpoofCeleb评估上以0.03%的CM-EER和0.038的最小a-DCF超越后者(后者分别为0.16%和0.047)。由于骨干网络共享且冻结,新分支可在不重新训练现有任务头和提示的情况下附加。在ASVspoof5上,面对对抗攻击和编解码失真,仅使用提供的数据即达到0.092的最小a-DCF和3.60%的CM-EER,表明该设计在现实野外威胁下依然稳健。

英文摘要

Spoofing-aware speaker verification (SASV) must confirm who is speaking and that the speech is genuine, but current systems gain one at the expense of the other. Modular cascades are accurate yet run two large self-supervised encoders, whereas end-to-end models are efficient but lose accuracy when a single embedding has to serve two conflicting objectives. We show that this trade-off between efficiency and specialization is avoidable. A single frozen W2V-BERT~2.0 backbone is adapted per task by deep hybrid Wavelet Prompt Tuning (WPT), so each branch obtains its own view of the shared encoder through dedicated prompts and a task head, and the scores are combined only at inference. Training only about 7M parameters, a small fraction of those a strong two-encoder cascade requires, the system surpasses it on SpoofCeleb evaluation with 0.03% CM-EER and 0.038 min a-DCF against 0.16% and 0.047. Because the backbone is shared and frozen, new branches attach without retraining the heads and prompts already in place. On ASVspoof5, with adversarial attacks and codec distortions, it reaches 0.092 min a-DCF and 3.60% CM-EER using only the provided data, indicating that the design holds under realistic in-the-wild threats.

Commentsaccepted at IEEE SLT 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑