arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于病因的对比严重度嵌入与音韵伪标注用于多语种构音障碍语音

Per-Aetiology Contrastive Severity Embeddings with Phonological Pseudo-Labelling for Multilingual Dysarthric Speech

Bernard Muller, Antonio Armando Ortiz Barrañón, LaVonne Roberts

arXiv 2609.21789首次发表:更新:

发表机构

The Scott-Morgan Foundation; Tecnológico de Monterrey; SMF Labs(斯科特·摩根基金会; 蒙特雷理工学院; SMF实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对多语种构音障碍严重度评估,提出基于病因的对比嵌入模型,结合音韵伪标注,在CP、PD和ALS上均优于混合基线,显著提升宏F1分数。

AI 中文摘要

大多数多语种构音障碍严重度系统要么在单一病因-语言对上训练,要么将异质病因合并到一个标签空间中。我们通过四个匹配的HuBERT-base对比嵌入模型,在共享骨干网络、训练方案、语料库注册表和留出评估下测试了这一合并假设:一个混合病因基线和三个针对脑性瘫痪(CP)、帕金森病(PD)和肌萎缩侧索硬化(ALS)的病因特异性模型。训练结合了临床标注语音和来自无训练音韵学剖析方法[1]、[2]的序数伪标签。在说话人不相交、泄漏过滤的留出子集上,病因特异性模型在所有三个目标病因上均优于混合基线:CP(宏F1 0.829对0.676,相对提升+22.6%)、PD(0.715对0.511,+40.0%)和ALS(0.788对0.596,+32.3%)。在CP上,添加144个SAP和44个CDSD伪标签说话人将宏F1从仅临床CP模型的0.786提升至0.829(+4.3个百分点)。每个病因的训练数据涵盖三到七种语言。我们将此定位为标签空间设计选择的受控比较,并讨论伪标签校准、分割卫生和置信度阈值部署作为未来工作的重要局限性。

英文摘要

Most multilingual dysarthria-severity systems either train on a single aetiology-language pair or pool heterogeneous aetiologies into one label space. We test that pooling assumption with four matched HuBERT-base contrastive embedding models under a shared backbone, training recipe, corpus registry and held-out evaluation: one mixed-aetiology baseline and three aetiology-specific models for cerebral palsy (CP), Parkinson's disease (PD) and amyotrophic lateral sclerosis (ALS). Training combines clinically labelled speech with ordinal pseudo-labels from a training-free phonological profiling method [1], [2]. On speaker-disjoint, leakage-filtered held-out subsets, the per-aetiology models outperform the mixed baseline across all three target aetiologies: CP (macro F1 0.829 vs 0.676, +22.6 % relative), PD (0.715 vs 0.511, +40.0 %) and ALS (0.788 vs 0.596, +32.3 %). On CP, adding 144 SAP and 44 CDSD pseudo-labelled speakers lifts macro F1 from 0.786 to 0.829 over a clinical-only CP model (+4.3 percentage points). Training data span three to seven languages per aetiology. We position this as a controlled comparison of label-space design choices and discuss pseudo-label calibration, split hygiene, and confidence-thresholded deployment as important limitations for future work.

CommentsAccepted at IEEE SLT 2026, 13-16 December 2026, Palermo, Sicily. v2: corrected pseudo-labelled subset language count from 'seven' to 'five' in Sec. 2 and Table II caption (arithmetic in the paper unchanged; the enumerated list already had five languages)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑