发表机构
Shanghai Jiao Tong University; City University of Hong Kong; The Hong Kong Polytechnic University; Beijing Foreign Studies University(上海交通大学; 香港城市大学; 香港理工大学; 北京外国语大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有语音强制对齐系统缺乏低资源语言变体专门模型的问题,利用17小时语料库和自定义词典,训练成都话依赖文本与独立文本的对齐器,显著优于标准普通话基线,建立实用自训练管道。
AI 中文摘要
语音强制对齐是语音研究中的关键技术,但现有对齐系统缺乏针对低资源语言变体的专门模型。我们使用一个17小时的语料库和一个自定义的G2P词典,为成都话训练了依赖文本和独立文本的对齐器。我们训练了一个依赖文本的GMM-HMM模型(Chengdu-MFA),并使用Chengdu-MFA的伪标签对预训练的音频编码器进行微调,用于独立文本对齐(Chengdu-FC)。在专家标注的测试集上的评估表明,这两种方法都显著优于标准普通话基线。Chengdu-MFA将平均音素边界差异降低了31.8%,而Chengdu-FC降低了61.2%。这项工作建立了一个实用的自训练管道,用于为资源不足的变体开发准确的对齐器,而无需耗费大量人力和时间的手动标注。
英文摘要
Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners for Chengdu Mandarin using a 17-hour corpus and a custom G2P dictionary. We trained a text-dependent GMM-HMM model (Chengdu-MFA) and fine-tuned a pretrained audio encoder on frame classification with Chengdu-MFA's pseudo label for text-independent alignment (Chengdu-FC). Evaluation on an expert-annotated test set show that both methods significantly outperform Standard Mandarin baselines. Chengdu-MFA reduced average phone boundary differences by 31.8%, while Chengdu-FC achieved a 61.2% reduction. This work establishes a practical bootstrapping pipeline for developing accurate aligners for under-resourced varieties without labor- and time-intensive manual annotation.
Comments5 pages, 1 figure