语音基础模型真的学习到词汇了吗?
Do speech foundation models really learn words?
- University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过残差化方法剔除音素信息,证明HuBERT和wav2vec 2.0在深层确实学习到独立于词形的词汇表示,并能增强词汇发现任务中的高阶语言信息。
AI中文摘要:
自监督语音基础模型现已被广泛应用于各种下游任务,包括传统语音识别,以及作为语音感知语言模型中词元的基础。理解其有用性的尝试主要集中在探究其表示区分音素和词汇的能力。然而,对词汇的区分能力并不必然意味着模型对词汇本身具有专门化的表示。良好的词汇区分能力可能源于对词形(音素)的良好编码,而非独立于词形的、编码词汇身份或句法/语义属性的词汇表示。通过使用残差化方法剔除音素信息,我们表明,在较深层中,HuBERT和wav2vec 2.0总体上确实学习到了能够以合理保真度编码词汇、且独立于局部语音内容的表示。我们进一步展示,这种简单的解耦方法能够增强词汇发现任务中的高阶语言信息。
英文摘要:
Self-supervised speech foundation models are now used in a wide array of downstream applications, including traditional speech recognition and as the basis for tokens in speech-aware language models. Attempts to understand their usefulness have largely focused on probing their representations' ability to discriminate phonemes and words. However, discriminative ability for words need not imply specialized representation of words per se. Good discrimination of words may be explained by good encoding of word form (phonemes) rather than form-independent word representations encoding identity or syntactic/semantic properties. By partialling out phoneme information using residualization, we show that, in later layers, HuBERT and wav2vec 2.0 do in general learn representations which encode words with reasonable fidelity independently of local phonetic content. We show that this simple approach to disentanglement can enhance higher-order linguistic information in word discovery tasks.