ToneCL:面向少样本音节级声调分类的对比学习
ToneCL: Contrastive Learning for Few-Shot Syllable-Level Tone Classification
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对低资源声调语言音节级标注困难,提出轻量级对比学习框架ToneCL,通过保持声调身份的增强预训练和少样本微调,在普通话和越南语上显著优于基线,跨语言迁移有效。
AI中文摘要:
声调语言占世界语言的50%-70%以上,但绝大多数是低资源语言,缺乏自动声调分类所需的大规模转录语料库。现有数据集通常在句子级别收集,而田野语言学家需要细粒度的音节级标注。我们提出ToneCL,一种用于少样本音节级声调分类的轻量级对比学习框架。我们在普通话和越南语上模拟低资源条件,将每个声调类别的标注数据限制为几十个示例。ToneCL在未标注语音上通过保持声调身份的增强进行预训练,然后在少样本示例上进行微调。实验表明,我们的方法持续优于基线,在10-shot的六说话人普通话任务上达到91.6%的准确率。跨语言迁移同样有效:在越南语上预训练并在普通话上微调,在10-shot时达到91.0%的准确率。消融实验证实,频带抑制是最关键的增强方法。
英文摘要:
Tone languages constitute over 50-70% of the world's languages, but the vast majority are low-resource, lacking the large transcribed corpora needed for automatic tone classification. Existing datasets are typically collected at the sentence level, whereas field linguists require fine-grained syllable-level annotations. We propose ToneCL, a lightweight contrastive learning framework for few-shot syllable-level tone classification. We simulate low-resource conditions on Mandarin and Vietnamese, limiting labeled data to tens of examples per tone class. ToneCL is pretrained on unlabeled speech with augmentations that preserve tonal identity, then fine-tuned on few-shot examples. Experiments show our method consistently outperforms baselines, achieving 91.6% on six-speaker Mandarin at 10 shots. Cross-lingual transfer is also effective: pretraining on Vietnamese and fine-tuning on Mandarin reaches 91.0\% accuracy at 10 shots. Ablation confirms that frequency band rejection is the most critical augmentation.