arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31293eess.AScs.CLcs.SD

为何阿尔茨海默病语音筛查难以泛化:通过跨语料证据锚定弥合部署差距

Why Alzheimer's Speech Screening Fails to Generalize: Bridging the Deployment Gap via Cross-Corpus Evidence Anchoring

  • Nanjing University of Posts and Telecommunications(南京邮电大学)
  • University of Science and Technology of China(中国科学技术大学)
  • Shouyi Technology(守一科技)
  • University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

Zijian Lu, Sizhe Liu, Yin Zhang, Jixuan Deng, Xinrong Lin, Xinchen Yuan, Chicheng Jin, Yiping Zuo, Yuanchao Li

AI总结:

本研究通过跨语料评估揭示阿尔茨海默病语音筛查模型泛化失败,提出融合证据锚定的方法,将最差域AUC从0.504提升至0.615。

AI中文摘要:

基于语音的筛查是检测阿尔茨海默病及相关认知风险的一种有前景的非侵入性方法。然而,在单一领域上训练的模型往往难以泛化到未见过的语言、任务或录音协议。本文通过跨四个不同数据集的留一语料库评估来研究这一部署差距。在70个可解释的语音和语言特征中,59个在健康对照组和认知风险组之间跨语料库表现出方向冲突,其中停顿、静默和语速显示出高度的协议敏感性。此外,虽然XLM-R文本基线取得了较强的平均性能,但其在表现最弱的留出域上的ROC曲线下面积(AUC)降至0.520。在相同协议下,标准GroupDRO基线达到0.766的平均说话人AUC和0.504的最差域AUC。为解决这一问题,我们提出了一种融合方法,将XLM-R文本基线分数与训练期间选择的证据锚点相结合。平衡融合实现了0.785的平均说话人AUC,而锚点重融合将最差情况下的说话人AUC提升至0.615。这项工作强调了在认知语音筛查中审计特征可迁移性并报告最差情况域鲁棒性的必要性。

英文摘要:

Speech-based screening is a promising, non-invasive approach for detecting Alzheimer's disease and related cognitive risks. However, models trained on a single domain often generalize poorly to unseen languages, tasks, or recording protocols. This paper investigates this deployment gap using a leave-one-corpus-out evaluation across four distinct datasets. Among 70 interpretable speech and language features, 59 exhibit direction conflicts between healthy control and cognitive risk groups across corpora, with pause, silence, and speech rate showing high protocol sensitivity. Furthermore, while the XLM-R text baseline achieves strong average performance, its Area Under the ROC Curve (AUC) drops to 0.520 on the weakest held-out domain. A standard GroupDRO baseline reaches a 0.766 mean speaker AUC and a 0.504 worst-domain AUC under the same protocol. To address this, we propose a fusion method that integrates XLM-R text baseline scores with evidence anchors selected during training. Balanced fusion achieves a 0.785 mean speaker AUC, while anchor-heavy fusion raises the worst-case speaker AUC to 0.615. This work highlights the need to audit feature transferability and report worst-case domain robustness in cognitive speech screening.

补充信息

↑