AI 中文总结
本研究针对真实课堂环境下的英语语音,适配WavLM-TDNN模型,对比ECAPA-TDNN,采用两阶段训练策略,实现等错误率显著降低,为教育类AI工具提供支撑。
AI 中文摘要
开发对课堂噪声具有鲁棒性且对儿童和成年说话人均有效的说话人验证(SV)模型,对支持教育环境的AI工具至关重要。本研究使用包含部分说话人身份标注的真实英语课堂数据集,大部分录音未标注。我们适配WavLM-TDNN模型用于课堂SV,与ECAPA-TDNN基线模型、在课堂数据上训练的ECAPA-TDNN模型相比,等错误率(EER)的平均相对降低分别达23.99%和6.32%。此外,我们研究了课堂SV的两种训练策略:自监督学习(SSL)和先以SSL预训练再用有限标注数据微调的两阶段方法。五折交叉验证表明,两阶段策略始终优于仅SSL方法,EER平均相对降低13.39%。
英文摘要
Developing speaker verification (SV) models that are robust to classroom noise and effective across both children and adult speakers is critical for AI tools supporting educational environments. In this study, we use a real-world English-speaking classrooms dataset containing partial speaker identity annotations, with most recordings remaining unlabeled. We adapt the WavLM-TDNN model for classroom SV, achieving average relative reductions in Equal Error Rate (EER) of 23.99% and 6.32% compared to the ECAPA-TDNN baseline and the ECAPA-TDNN model trained on classroom data, respectively. Additionally, we investigate two training strategies for SV in classroom settings: self-supervised learning (SSL) and a two-stage approach that first pre-trains with SSL and then fine-tunes with limited annotated data. Five-fold cross-validation demonstrates that the two-stage strategy consistently outperforms the SSL-only approach, achieving an average relative EER reduction of 13.39%.