发表机构
Graduate School of Informatics, Kyoto University; NTT, Inc.(京都大学信息学研究生院; 日本电信电话公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对语音到国际音标转录中标签不准确问题进行零样本语音分类评估,提出基于连续发音特征向量的分类方法,该方法优于离散标记法,且采用最佳时间聚合可提升分类效果,尤其对稀有音素和特定语音分类。
AI 中文摘要
近期用于语音到国际音标转录的语音基础模型依赖音素到音标的标签,但这些标签在语音上不一定准确。为研究此问题,我们对汉语送气音和日语 moraic 鼻音进行零样本语音分类评估。在排除这两种语言的音素到音标标签数据上训练的模型在两项任务中准确率低,表明离散国际音标标记的多语言覆盖对未见设置不足。为克服此限制,我们提出基于从各帧提取的连续发音特征向量的分类方法。该方法优于基于离散标记的方法,尤其是对稀有音素。我们还表明采用发音特征向量的最佳时间聚合对目标区分至关重要:单帧分类对送气音最佳,而分段分类显著改善鼻音分类。
英文摘要
Recent Phonetic Foundation Models (PFMs) for Speech-to-IPA transcription rely on Grapheme-to-Phoneme (G2P) labels, but the phoneme labels are not necessarily phonetically faithful. To investigate this issue, we evaluate zero-shot phonetic classification on Chinese aspiration and Japanese moraic nasals. A PFM trained on G2P-labeled data excluding these two languages yields poor accuracy on both tasks, showing that multilingual coverage with discrete IPA tokens is not sufficient for unseen settings. To overcome this limitation, we propose a classification method based on continuous Articulatory Feature (AF) vectors extracted from each frame. This AF-based approach outperforms discrete token-based methods, particularly for rare phones. We further show that it is crucial to adopt the optimal temporal aggregation of AF vectors for the target distinction: single-frame classification is best for aspiration, while segmental classification substantially improves nasal classification.
CommentsAccepted at Interspeech 2026