适用于跨声调与非声调语言的数据高效音素级多语种自动语音识别的潜在Softmax方法
Latent Softmax for Data-Efficient Phoneme-Based Multilingual ASR Across Tonal and Non-Tonal Languages
浏览论文内容
中文总结 AI 辅助
该研究针对多语种ASR中声调与非声调语言监督粒度不匹配的问题,提出Latent Softmax方法,在多语种语音识别实验中显著降低音素与混合错误率,提升了跨语言语音识别性能。
中文摘要 AI 辅助
基于音素的多语种自动语音识别(ASR)相比特定语言的子词建模,能更直接地在不同语言间共享声学证据。然而,当联合训练声调语言与非声调语言时,二者的监督粒度并不匹配:声调语言标注带调元音,而非声调语言通常仅提供基础元音标签。标准Softmax要么将二者视为不相关类别,削弱跨语言共享;要么合并声调,丢失声调语言所需的区分信息。我们提出Latent Softmax,一种兼容连接时序分类(CTC)的输出层,将带调元音建模为子类、基础元音建模为主类,辅音与CTC空白则保持单例标签。当仅观测到基础元音主类标签时,带调元音子类被视为潜在变量并进行边缘化处理。在AISHELL-1普通话与LibriSpeech英语上开展的多语种实验显示,Latent Softmax相比标准Softmax多语种基线,在AISHELL-1上将音素错误率(S2P)降低8.4%,在LibriSpeech test-clean上降低17.5%,在test-other上降低12.6%。改进后的音素编码器还为大语言模型音素到 grapheme转换及基于投影器的接口带来一致的词错误率提升。经代码切换适配后,Latent Softmax进一步在ASRU2019上将基于投影器的混合错误率降低2.6%,在CS-Dialogue数据集上降低9.5%。
英文摘要
Phoneme-based multilingual automatic speech recognition (ASR) can share acoustic evidence across languages more directly than language-specific subword modeling. When tonal and non-tonal languages are jointly trained, however, their supervision granularity does not match: tonal languages annotate tone-marked vowels, whereas non-tonal languages typically provide only base-vowel labels. A standard softmax either treats the two as unrelated classes, weakening cross-lingual sharing, or collapses tones, losing distinctions required by tonal languages. We propose Latent Softmax, a connectionist temporal classification (CTC)-compatible output layer that models tone-marked vowels as subclasses and base vowels as major classes, while consonants and the CTC blank remain singleton labels. When only a base-vowel major-class label is observed, the tone-marked vowel subclass is treated as latent and marginalized out. Multilingual experiments on AISHELL-1 Mandarin and LibriSpeech English show that Latent Softmax reduces speech-to-phoneme (S2P) phoneme error rates over a standard softmax multilingual baseline by 8.4% on AISHELL-1, 17.5% on LibriSpeech test-clean, and 12.6% on test-other. The improved speech-to-phoneme encoders also yield consistent word error rate gains for both large-language-model phoneme-to-grapheme conversion and projector-based interfaces. After code-switching adaptation, Latent Softmax further reduces projector-based mixed error rate by 2.6% on ASRU2019 and 9.5% on CS-Dialogue datasets.
发表机构
- TasiTech(塔西科技)
机构由 AI 辅助整理,请以论文原文为准。