arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

音节级别的无监督语音识别

Unsupervised Speech Recognition at the Syllable Level

Liming Wang, Kai-Wei Chang, Kunio Kashino, David Harwath, Mark Hasegawa-Johnson, James R. Glass

arXiv 2608.22907首次发表:更新:

AI 中文总结

该研究针对现有无监督语音识别方法依赖G2P且训练不稳定的问题,提出基于掩码语言建模的音节级UASR框架,在LibriSpeech上使CER相对降低40%,并有效推广到低资源语言。

AI 中文摘要

利用未配对的语音和文本训练语音识别器,即无监督语音识别(UASR),是将自动语音识别(ASR)扩展到长尾分布的低资源语言,并从非平行数据中实现多模态学习的关键一步。然而,现有的基于音素的方法通常依赖于 grapheme-to-phoneme 转换器(G2P)等成本高昂的资源,且由于训练不稳定,难以推广到音素边界模糊的语言。在本文中,我们通过引入一种基于掩码语言建模的音节级 UASR 框架来解决这两个挑战,该框架无需 G2P,也避免了基于生成对抗网络(GAN)方法的不稳定性。我们的方法在 LibriSpeech 上实现了字符错误率(CER)最高达 40% 的相对降低,并且能有效推广到以往方法难以处理的低资源语言。代码已公开。

英文摘要

Training speech recognizers with unpaired speech and text -- known as unsupervised speech recognition (UASR) -- is a crucial step toward extending ASR to low-resource languages in the long-tail distribution and enabling multimodal learning from non-parallel data. However, existing approaches based on phones often rely on costly resources such as grapheme-to-phoneme converters (G2Ps) and struggle to generalize to languages with ambiguous phoneme boundaries due to training instability. In this paper, we address both challenges by introducing a syllable-level UASR framework based on masked language modeling, which avoids the need for G2P and the instability of GAN-based methods. Our approach achieves up to a 40\% relative reduction in character error rate (CER) on LibriSpeech and generalizes effectively to low-resource languages that have remained particularly difficult for prior methods. Code is publicly available\footnote{https://github.com/cactuswiththoughts/SylCipher}.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑