发表机构
Texas A&M University(德克萨斯A&M大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AccentCL提出一种类增量学习框架,利用冻结编码器和不平衡感知损失,实现英语口音分类的鲁棒性,并支持新口音类别的增量添加,无需完全重训。
AI 中文摘要
口音分类器通常使用固定的标签清单进行训练,当新数据可用时,无法容纳新的口音类别。此外,带口音的语音语料库通常表现出显著的类别不平衡和/或领域偏移,这是由于不同语料库之间的录音条件差异所致。我们提出了AccentCL,一个用于英语口音分类的类增量学习框架,该框架对类别不平衡和跨语料库领域偏移具有鲁棒性。AccentCL从冻结的Whisper-Large-v3编码器中提取多层表示,并通过不平衡感知的交叉熵损失进行优化,以减少对多数口音类别的偏见,同时使用领域均值对齐损失来最小化训练语料库之间的分布均值偏移。随后,通过基于回放的持续学习扩展标签空间,使用冻结的基础模型进行知识保留,并采用旧到新边际损失来减少对新添加类别的过度预测。在五类口音分类任务中,AccentCL实现了77.1%的平衡准确率和76.9%的宏平均F1分数。我们进一步评估了模型增量整合两个新口音类别(西班牙口音英语和中文口音英语)的能力。当将西班牙口音英语添加到预训练模型时,AccentCL在新类别上达到83.3%的F1分数,同时在基础类别上保持77.3%的平衡准确率。随后添加中文口音英语时,它在保留先前学习类别77.6%平衡准确率的同时,在新类别上实现了61.8%的F1分数。这些结果表明,AccentCL能够实现鲁棒的地区口音分类,同时允许在不进行完全重新训练的情况下添加新的口音类别。
英文摘要
Accent classifiers are typically trained with a fixed label inventory and cannot accommodate new accent categories as new data becomes available. Moreover, accented speech corpora often exhibit substantial class imbalance and/or domain shift due to differences in recording conditions across corpora. We present AccentCL, a class-incremental learning framework for English accent classification that is robust to class imbalance and cross-corpus domain shift. AccentCL extracts multi-layer representations from a frozen Whisper-Large-v3 encoder, optimized with an imbalance-aware cross-entropy loss to reduce bias toward the majority accent classes and a domain mean alignment loss that minimizes distributional mean shift across training corpora. The label space is then expanded via replay-based continual learning, using the frozen base model for knowledge retention and an old-to-new margin loss to reduce overprediction on newly added classes. On a five-class accent classification task, AccentCL achieves 77.1% balanced accuracy and a 76.9% macro-averaged F1 score. We further evaluate the model's ability to incrementally incorporate two new accent categories: Spanish-accented and Chinese-accented English. When adding Spanish-accented English to the pretrained model, AccentCL attains an F1 of 83.3% on the new class while retaining 77.3% balanced accuracy on the base classes. When subsequently adding Chinese-accented English, it achieves 61.8% F1 on the new class while preserving 77.6% balanced accuracy on the previously learned classes. These results show that AccentCL enables robust regional accent classification while allowing new accent categories to be added without full retraining.
CommentsPublished in Proceedings of IEEE Spoken Language Technology Workshop (SLT) 2026