Whisper-PMFA:使用 Whisper 模型进行说话人验证的部分多尺度特征聚合
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
- Tsinghua University(清华大学)
- Shenzhen Research Institute of Big Data(深圳大数据研究院)
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出基于 Whisper 编码器块子集的部分多尺度特征聚合(PMFA)方法用于说话人验证,在 VoxCeleb1 和 CN-Celeb1 上取得优异性能,并验证了多语言训练的跨语言鲁棒性及低秩适应的高效性。
AI中文摘要:
本文提出将用于自动语音识别的大规模预训练模型 Whisper 应用于说话人验证。基于 Whisper 编码器块的子集,提出了一种部分多尺度特征聚合(PMFA)方法,以提取高判别性的说话人嵌入。实验结果表明,使用 Whisper 编码器中后部的块能保留更多说话人信息。在 VoxCeleb1 和 CN-Celeb1 数据集上,本系统分别实现了 1.42% 和 8.23% 的等错误率(EER),相比 ECAPA-TDNN 基线绝对 EER 降低了 0.58% 和 1.81%,相比 ResNet34 基线降低了 0.46% 和 0.97%。此外,结果表明使用在多语言数据上训练的 Whisper 模型能有效增强模型跨语言的鲁棒性。最后,评估了低秩适应方法,该方法将可训练模型参数减少了约 45 倍,同时 EER 仅略微增加 0.2%。
英文摘要:
In this paper, Whisper, a large-scale pre-trained model for automatic speech recognition, is proposed to apply to speaker verification. A partial multi-scale feature aggregation (PMFA) approach is proposed based on a subset of Whisper encoder blocks to derive highly discriminative speaker embeddings.Experimental results demonstrate that using the middle to later blocks of the Whisper encoder keeps more speaker information. On the VoxCeleb1 and CN-Celeb1 datasets, our system achieves 1.42% and 8.23% equal error rates (EERs) respectively, receiving 0.58% and 1.81% absolute EER reductions over the ECAPA-TDNN baseline, and 0.46% and 0.97% over the ResNet34 baseline. Furthermore, our results indicate that using Whisper models trained on multilingual data can effectively enhance the model's robustness across languages. Finally, the low-rank adaptation approach is evaluated, which reduces the trainable model parameters by approximately 45 times while only slightly increasing EER by 0.2%.