通过声学掩蔽量化辅音对单词可懂度的贡献
Quantifying Consonant Contributions to Word Intelligibility via Acoustic Masking
- University of Texas at Austin(德克萨斯大学奥斯汀分校)
- Arizona State University(亚利桑那州立大学)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出基于声学掩蔽的可扩展方法,通过自动语音识别模型计算掩蔽诱导误识别率来量化辅音对单词可懂度的贡献,并验证其与音素频率和功能负荷的相关性,发现辅音贡献具有语言依赖性。
AI中文摘要:
辅音对单词是否被理解所起的作用并不相同。鉴于治疗可用时间有限,根据辅音对可懂度的贡献对其进行排序,有助于在运动性言语障碍中确定干预目标的优先级。然而,测量这种贡献依赖于难以扩展的感知研究。本文提出了一种可扩展的方法,利用声学掩蔽来测量辅音贡献。我们在孤立词中一次静音一个辅音,并测试自动语音识别(ASR)模型是否仍能识别该单词。我们将辅音的贡献分数定义为其被掩蔽实例中单词被误识别的比例,称之为掩蔽诱导误识别率(MMR)。我们根据先前报道与辅音贡献相关的两个语言因素,即音素频率和功能负荷,对MMR进行了验证。我们使用三种ASR架构——MMS(仅编码器)、Whisper(编码器-解码器)和Qwen3-ASR(基于大语言模型)——在英语、西班牙语、德语和捷克语四种语言中进行了这一分析。通过偏Spearman相关分析,我们发现音素频率与MMR呈负相关,而功能负荷与MMR呈正相关。换言之,更频繁的辅音在被掩蔽时破坏性较小,而承载更多词汇对比的辅音则更具破坏性。进一步的跨语言分析表明,辅音排序并不一致,表明辅音贡献具有语言依赖性。
英文摘要:
Consonants contribute unequally to whether a word is understood. Given the limited time available for therapy, ranking consonants by contribution to intelligibility helps prioritize intervention targets in motor speech disorders. However, measuring this contribution relies on perceptual studies that are difficult to scale. This paper presents a scalable method that measures consonant contribution using acoustic masking. We silence one consonant at a time in an isolated word and test whether an automatic speech recognition (ASR) model still recognizes the word. We define a consonant's contribution score as the proportion of its masked instances for which the word becomes misrecognized, which we refer to as the mask-induced misrecognition rate (MMR). We relate MMR to two linguistic factors previously reported to correlate with consonant contribution, namely phoneme frequency and functional load. We apply this analysis across four languages, English, Spanish, German, and Czech, using three ASR architectures, MMS (encoder-only), Whisper (encoder-decoder), and Qwen3-ASR (LLM-based). Using partial Spearman correlations, we find that phoneme frequency correlates negatively with MMR while functional load correlates positively. In other words, more frequent consonants are less disruptive when masked, whereas consonants carrying more lexical contrast are more disruptive. Further cross-language analysis shows that consonant rankings agree only partially across languages, indicating that consonant contribution is language-dependent.