用于低资源澳大利亚原住民语言识别的混合持续学习
Hybrid Continual Learning for Low-Resource Australian Aboriginal Language Identification
- University of New South Wales(新南威尔士大学)
- University of Melbourne(墨尔本大学)
- Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对低资源澳大利亚原住民语言识别中数据稀缺和灾难性遗忘问题,提出重放增强弹性权重巩固和约束引导知识蒸馏两种混合持续学习方法,实验表明其优于微调及现有基线,能提升对多种AAL的适应性并保持对高资源语言的性能。
AI中文摘要:
语言识别对于将濒危的澳大利亚原住民语言(AALs)整合到支持语言振兴和数字包容的语音技术中至关重要。然而,极端的数据稀缺限制了模型性能。从高资源语言进行迁移学习有前景,但适应新语言时常遭受灾难性遗忘。持续学习(CL)可缓解此问题,但在数据非常有限时仍具挑战。为此,我们提出两种混合持续学习方法:重放增强弹性权重巩固和约束引导知识蒸馏,以在保留先前所学知识的同时,使预训练语音模型适应AAL识别。在瓦尔皮里语、达拉邦语和达拉瓦尔语上的实验表明,所提方法优于微调及现有CL基线,在保持对先前学习的高资源语言性能的同时,提升了对多种AAL的适应性。
英文摘要:
Language identification is an important step toward integrating endangered Australian Aboriginal languages (AALs) into speech technologies supporting language revitalisation and digital inclusion. However, extreme data scarcity limits model performance. Transfer learning from high-resource languages shows promise but often suffers from catastrophic forgetting when adapting to new languages. Continual learning (CL) can mitigate this issue, though it remains challenging with very limited data. To address this, we propose two hybrid continual learning methods: Replay Augmented Elastic Weight Consolidation and Constraint Guided Knowledge Distillation to adapt pretrained speech models for AAL identification while preserving previously learned knowledge. Experiments on Warlpiri, Dalabon and Dharawal show that the proposed methods outperform fine-tuning and existing CL baselines, improving adaptation to multiple AALs while maintaining performance on previously learnt high-resource languages.