arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13694eess.AScs.SD

基于最优传输的次音素声学建模用于发音评估

Subphonetic Acoustic Modeling via Optimal Transport for Pronunciation Assessment

  • The University of Tokyo(东京大学)
  • Chunghwa Telecom Co., Ltd.(中华电信股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Haopeng Geng, Jiun-Ting Li, Daisuke Saito, Nobuaki Minematsu

AI总结:

本文提出一种拓扑感知的逐帧声学模型,结合有序次音素状态与最优时间传输分类,实现密集单调的状态判别,提升发音评估中的分割精度与错音检测性能。

AI中文摘要:

发音评估需要时间上精确、具有诊断意义且忠实于学习者实际发音的声学证据。然而,现有的声学模型往往难以同时提供识别和分割证据。基于CTC的音素识别器可以灵活地预测音素序列,但其稀疏且尖峰的后验概率常常遗漏音素边界和细粒度的发音线索。相比之下,当转录文本可用时,文本相关的强制对齐器能提供可靠的时间信息,但无法直接应用于无参考的发音分析。在这项工作中,我们提出了一种拓扑感知的逐帧声学模型,该模型在每个音素内学习密集的有序状态后验概率。关键思想是通过将有序的次音素状态与最优时间传输分类(OTTC)相结合,在神经声学模型中恢复音素内部的状态结构。这种组合鼓励密集的单调逐帧状态判别,同时保持音素识别能力。在朗读、自发和二语语音上的实验表明,与神经基线相比,分割性能有所提升,且识别性能具有竞争力。下游评估进一步展示了在错音检测和自动发音评估方面的增益。探针分析表明,学习到的状态捕获了依赖于音素的声学结构,而非任意的逐帧分布。

英文摘要:

Pronunciation assessment requires acoustic evidence that is temporally precise, diagnostically meaningful, and faithful to the learner's actual production. However, existing acoustic models often struggle to provide recognition and segmentation evidence simultaneously. CTC-based phone recognizers can predict phone sequences flexibly, but their sparse and peaky posteriors often miss phone boundaries and fine-grained pronunciation cues. In contrast, text-dependent forced aligners provide reliable temporal information when transcripts are available, but are not directly applicable to reference-free pronunciation analysis. In this work, we propose a topology-aware frame-wise acoustic model that learns dense ordered state posteriors within each phone. The key idea is to recover phone-internal state structure in a neural acoustic model by combining ordered subphonetic states with optimal temporal transport classification (OTTC). This combination encourages dense monotonic frame-level state discrimination while preserving phone recognition ability. Experiments on read, spontaneous, and L2 speech show improved segmentation over neural baselines with competitive recognition performance. Downstream evaluations further show gains in mispronunciation detection and automatic pronunciation assessment. Probing analysis suggests that the learned states capture phoneme-dependent acoustic structure rather than arbitrary frame-level distributions.

补充信息

↑