arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13425cs.CLeess.ASeess.SP

运动、认知还是语料?基于语音的帕金森病检测中跨语言迁移能保留什么

Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection

Serli Kopar, Sam Gijsen, Abner Hernandez, Paula Andrea Perez-Toro, Kerstin Ritter

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过对9个SSL语音骨干的分层分析,探究跨语言迁移中保留的语音特征类型,发现层选择依赖语料库且迁移信号缺乏病理特异性,明确了基于语音的PD检测模型的关键局限。

中文摘要 AI 辅助

自监督学习(SSL)语音表示在单个语料库内的帕金森病(PD)检测中表现出强劲性能,但目前仍不清楚这些模型是捕捉了疾病相关特征还是利用了数据集特有的混杂因素,尤其是大多数SSL骨干仅在健康语音上进行预训练。为探究该问题,我们在三种语言上使用低容量逻辑回归探针对9个SSL语音骨干进行分层分析,将评估构建为逐步引入参与者身份、录音条件、语言和病理分布偏移的多种场景。结果揭示两项关键发现:第一,层选择高度依赖语料库,最优表示层主要由源数据集而非SSL架构本身决定;第二,迁移的判别信号缺乏病理特异性,在目标语料库中,训练用于检测PD的分类器对PD和痴呆语音分配的概率同样高。这些结果凸显了基于语音的病理识别模型在临床环境中可靠部署前必须解决的关键局限性。

英文摘要

Self-supervised learning (SSL) speech representations achieve strong performance for Parkinson's disease (PD) detection within individual corpora. However, it remains unclear whether these models capture disease-related characteristics or exploit dataset-specific confounds, particularly since most SSL backbones are pretrained exclusively on healthy speech. To investigate this question, we perform a layer-wise analysis of nine SSL speech backbones using a low-capacity logistic regression probe across three languages. We structure the evaluation as multiple scenarios that progressively introduce distribution shifts in participant identity, recording conditions, language, and pathology. Our results reveal two key findings. First, layer selection is highly corpus-dependent: the optimal representation layer is determined primarily by the source dataset rather than by the SSL architecture itself. Second, the transferred discriminative signal lacks pathological specificity: classifiers trained to detect PD assign similarly high probabilities to both PD and dementia speech in the target corpus. These results highlight critical limitations that must be addressed before speech-based pathology recognition models can be reliably deployed in clinical settings.

发表机构

  • Hertie Institute for AI in Brain Health, University of Tübingen(蒂宾根大学赫蒂脑健康人工智能研究所)
  • Tübingen AI Center, University of Tübingen(蒂宾根大学蒂宾根人工智能中心)
  • Charité–Universitätsmedizin(夏里特医学院)
  • Friedrich-Alexander-Universität Erlangen-Nürnberg(埃尔朗根-纽伦堡弗里德里希-亚历山大大学)

机构由 AI 辅助整理,请以论文原文为准。

↑