arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于发音特征分解的多任务非标准音素识别

Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition

Sophia Riaz, Haoze Zheng, Amos Roche, Miyu Zhang, Anamika Ragu, Salvatore Penachio, Kaustav Mukherjee, Aneesh Jonelagadda

arXiv 2608.22273首次发表:更新:

发表机构

Kaliber AI(卡利伯人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出分层多任务学习结合交叉注意力融合模块与半监督学习的方法,在L2-ARCTIC数据集上显著提升非标准音素识别性能,为病理语音等非标准语音的鲁棒可解释识别提供新策略。

AI 中文摘要

病理语音及更广泛的非标准语音因与标准发音存在系统性偏差,且带标注的临床语音数据有限,给自动音素识别带来重大挑战。现有音素识别系统通常在标准语音上训练,并将音素视为原子类别标签,限制了其检测语音障碍和口音中常见结构化发音错误的能力。本研究提出一种语言学结构化的非标准音素识别方法,将音素预测分解为方式、部位、浊音等发音特征维度。我们采用分层多任务学习架构实现该方案,其中特定任务的发音特征头学习特征级表示,随后通过基于交叉注意力的融合模块整合以生成音素预测。为解决病理语音标注的稀缺性与噪声问题,我们将该框架与基于动量伪标签(Momentum Pseudo-Labeling,MPL)的半监督学习相结合,并提出级联训练策略,逐步引入发音特征任务,同时分阶段解冻预训练语音编码器。在用作病理语音变异代理的L2-ARCTIC数据集上的实验表明,与强基线架构相比,所提方法在音素识别性能上实现了显著提升,同时产生与音系特征结构一致的可解释错误模式。这些结果表明,发音特征监督是针对非标准语音实现鲁棒且可解释的音素识别的有前景策略,并为未来在临床诊断的病理语音数据集上的验证提供了动力。

英文摘要

Pathological and more broadly non-canonical speech present significant challenges for automatic phoneme recognition due to systematic deviations from canonical pronunciation and limited availability of labeled clinical speech data. Existing phoneme recognition systems are typically trained on canonical speech and treat phonemes as atomic categorical labels, limiting their ability to detect structured articulatory errors common in speech disorders and accents. In this work, we introduce a linguistically structured approach to non-canonical phoneme recognition that decomposes phoneme prediction into articulatory feature dimensions such as manner, place, and voicing. We implement this formulation using a hierarchical multi-task learning architecture in which task-specific articulatory feature heads learn feature-level representations that are subsequently integrated through a cross-attention-based fusion module to produce phoneme predictions. To address the scarcity and noise of pathological speech labels, we combine this framework with semi-supervised learning via Momentum Pseudo-Labeling (MPL) and propose a cascaded training strategy that progressively introduces articulatory feature tasks while employing staged unfreezing of a pretrained speech encoder. Experiments on L2-ARCTIC, used as a proxy for pathological speech variation, show that the proposed approach achieves substantial improvements in phoneme recognition performance compared to strong baseline architectures, while yielding interpretable error patterns aligned with phonological feature structure. These results suggest that articulatory feature supervision is a promising strategy for robust and interpretable phoneme recognition in non-canonical speech, and motivate future validation on clinically diagnosed pathological speech datasets.

Comments23 pages, 4 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑