发表机构
CMU; UT Austin; UTokyo; UBC(卡内基梅隆大学; 德克萨斯大学奥斯汀分校; 东京大学; 不列颠哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究音素分割与识别,利用基于S3M的音系激活映射(SPAM),引入识别和分割两个轻量级预测头,只需少量语音转录,能推广到未见音素,在多数据集上取得强分割和识别性能。
AI 中文摘要
音素分割和识别本质上是相关任务,但现代方法通常分别对它们进行建模。我们认为语音结构已潜藏在自监督语音模型(S3M)的表示中,只需引导其解决这两个任务。我们利用基于S3M的音系激活映射(SPAM),将每个S3M表示帧映射到音系特征激活向量。在此之上,引入两个简单有效的轻量级、无梯度下降预测头:识别头和分割头。该方法只需不到一分钟的语音转录,且能推广到训练中未见过的音素。在各种数据集上都取得了很强的分割和识别性能。
英文摘要
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.
CommentsCode will be released after acceptance