发表机构
University of Southern California; The University of Texas at Austin; Wonkwang University(南加州大学; 德克萨斯大学奥斯汀分校; 全北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出语言正交化方法,通过闭式岭残差化去除自监督语音表示中的语言信息,提升跨语言帕金森病检测性能,在多个骨干和语言上验证有效。
AI 中文摘要
自监督语音模型(S3Ms)为帕金森病(PD)检测提供了强大的表示,使得跨语言迁移对缺乏标注患者语音的语言具有吸引力。然而,这些表示也编码了语言身份,这可能混淆这种迁移:在没有目标语言PD语音的情况下,分类器可能区分语言而非病理,导致对目标患者产生高特异性但低敏感性。我们提出语言正交化,一种对S3M特征相对于外部VoxLingua107语言嵌入的闭式岭残差化方法,仅使用健康对照(HC)语音进行拟合。通过移除可预测语言的成分同时保留与病理相关的变异,它产生了一种较少依赖语言的几何结构,其中HC表示集中而PD表示分散。在五个S3M骨干网络、三个语音任务和三种目标语言上,我们的方法一致地提高了跨语言PD检测性能,同时纠正了高特异性/低敏感性的失败。
英文摘要
Self-supervised speech models (S3Ms) provide powerful representations for Parkinson's disease (PD) detection, making cross-lingual transfer attractive for languages lacking labeled patient speech. However, these representations also encode language identity, which can confound this transfer: without target-language PD speech, classifiers may separate languages rather than pathology, yielding high specificity but low sensitivity on target patients. We propose \emph{language orthogonalization}, a closed-form ridge residualization of S3M features against external VoxLingua107 language embeddings, fitted using only healthy-control (HC) speech. By removing language-predictable components while retaining pathology-related variation, it produces a less language-dependent geometry in which HC representations concentrate while PD representations disperse. Across five S3M backbones, three speech tasks, and three target languages, our method consistently improves cross-lingual PD-detection performance while correcting the high-specificity/low-sensitivity failure.
CommentsIEEE SLT 2026 submission