arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过音系激活映射进行音素分割与识别

Phone Segmentation and Recognition through Phonological Activation Mapping

Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh, Chin-Jou Li, Eunjung Yeo, Daisuke Saito, Nobuaki Minematsu, Shinji Watanabe, Jian Zhu, David Harwath, David R. Mortensen

arXiv 2607.09020首次发表:更新:

发表机构

CMU; UT Austin; UTokyo; UBC(卡内基梅隆大学; 德克萨斯大学奥斯汀分校; 东京大学; 不列颠哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究音素分割与识别,利用基于S3M的音系激活映射(SPAM),引入识别和分割两个轻量级预测头,只需少量语音转录,能推广到未见音素,在多数据集上取得强分割和识别性能。

AI 中文摘要

音素分割和识别本质上是相关任务,但现代方法通常分别对它们进行建模。我们认为语音结构已潜藏在自监督语音模型(S3M)的表示中,只需引导其解决这两个任务。我们利用基于S3M的音系激活映射(SPAM),将每个S3M表示帧映射到音系特征激活向量。在此之上,引入两个简单有效的轻量级、无梯度下降预测头:识别头和分割头。该方法只需不到一分钟的语音转录,且能推广到训练中未见过的音素。在各种数据集上都取得了很强的分割和识别性能。

英文摘要

Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.

CommentsCode will be released after acceptance

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑