可微且严重度不变的离散令牌用于构音障碍语音识别
Differentiable and Severity-invariant Discrete Tokens for Dysarthric Speech Recognition
查看机构详情
- The Chinese University of Hong Kong(香港中文大学)
- Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
- National Research Council Canada(加拿大国家研究委员会)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出可微且严重度不变的离散令牌方法,与构障碍语音识别任务紧密集成,在UASpeech和TORGO上显著降低词错误率,并通过正则化减少严重度组间差异。
中文摘要 AI 辅助
本文提出了新颖的可微且严重度不变(DSI)离散令牌方法,这些方法不仅与下游构音障碍语音识别任务紧密集成,而且最小化了不同语音损伤严重度组之间的离散令牌多样性。在UASpeech和TORGO语料库上进行的实验表明,使用DSI令牌训练的Conformer模型在两项任务上分别优于可比较的基线HuBERT离散/连续特征,实现了统计显著的词错误率(WER)绝对降低2.22%/0.78%(相对降低9.14%/3.41%)和绝对降低1.78%/1.06%(相对降低18.43%/11.86%)。经过系统组合后,在UASpeech和TORGO上分别获得了最低的WER 18.90%和6.38%。音素特定的T-SNE可视化显示,严重度不变的正则化通过产生更大的重叠和更不明显的严重度组分布边界,减少了依赖严重度的变异性。
英文摘要
This paper proposes novel differentiable and severity-invariant (DSI) discrete token approaches that are not only tightly integrated with downstream dysarthric speech recognition tasks, but also minimise discrete token diversity across speech impairment severity groups. Experiments conducted on the UASpeech and TORGO corpora suggest that Conformer models trained using the DSI tokens outperform the comparable baseline HuBERT discrete/continuous features by statistically significant WER reductions of 2.22\%/0.78\% absolute (9.14\%/3.41\% relative) and 1.78\%/1.06\% absolute (18.43\%/11.86\% relative) on the two tasks, respectively. After system combination, the lowest WERs of 18.90\% and 6.38\% were obtained on UASpeech and TORGO. Phoneme-specific T-SNE visualizations show that severity-invariant regularization reduces severity-dependent variation by producing greater overlap and less distinct boundaries among severity-group distributions.