AI 中文总结
该研究针对现有无监督语音识别方法依赖G2P且训练不稳定的问题,提出基于掩码语言建模的音节级UASR框架,在LibriSpeech上使CER相对降低40%,并有效推广到低资源语言。
AI 中文摘要
利用未配对的语音和文本训练语音识别器,即无监督语音识别(UASR),是将自动语音识别(ASR)扩展到长尾分布的低资源语言,并从非平行数据中实现多模态学习的关键一步。然而,现有的基于音素的方法通常依赖于 grapheme-to-phoneme 转换器(G2P)等成本高昂的资源,且由于训练不稳定,难以推广到音素边界模糊的语言。在本文中,我们通过引入一种基于掩码语言建模的音节级 UASR 框架来解决这两个挑战,该框架无需 G2P,也避免了基于生成对抗网络(GAN)方法的不稳定性。我们的方法在 LibriSpeech 上实现了字符错误率(CER)最高达 40% 的相对降低,并且能有效推广到以往方法难以处理的低资源语言。代码已公开。
英文摘要
Training speech recognizers with unpaired speech and text -- known as unsupervised speech recognition (UASR) -- is a crucial step toward extending ASR to low-resource languages in the long-tail distribution and enabling multimodal learning from non-parallel data. However, existing approaches based on phones often rely on costly resources such as grapheme-to-phoneme converters (G2Ps) and struggle to generalize to languages with ambiguous phoneme boundaries due to training instability. In this paper, we address both challenges by introducing a syllable-level UASR framework based on masked language modeling, which avoids the need for G2P and the instability of GAN-based methods. Our approach achieves up to a 40\% relative reduction in character error rate (CER) on LibriSpeech and generalizes effectively to low-resource languages that have remained particularly difficult for prior methods. Code is publicly available\footnote{https://github.com/cactuswiththoughts/SylCipher}.