发表机构
Hertie Institute for AI in Brain Health; Tübingen AI Center; Sony Group Corporation; Friedrich-Alexander-University Erlangen-Nürnberg(赫尔蒂人工智能脑健康研究所; 图宾根人工智能中心; 索尼集团公司; 弗里德里希-亚历山大大学埃尔朗根-纽伦堡)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究声学因素能否系统影响基于语音的阿尔茨海默病评估预测,通过干预实验发现噪声等可改变预测,提出干预鲁棒性测试应成为临床语音模型标准。
AI 中文摘要
基于语音的阿尔茨海默病(AD)评估日益依赖于预训练的自监督学习(SSL)模型,这些模型直接从原始音频中学习声学表示,使模型暴露于录音因素。我们探究这些因素是否仅仅被编码在SSL表示中,还是能够系统地改变预测结果。使用ADReSSo数据集和三个大型SSL骨干网络,我们对仅参与者语音、非语音和完整录音音频施加受控的噪声和混响干预。我们结合逐层线性解码、输入和表示空间干预以及几何对齐分析,以区分声学可解码性与对AD预测的影响。我们的结果表明,受控的声学干预在所有三个SSL骨干网络中均改变了AD预测。噪声,尽管在原始数据中未显示显著的诊断组差异,却产生了最强的干预效果。重要的是,这些效果相对于分类器的决策方向是系统性地结构化的,在保留的测试集上可复现,并且当表示空间干预方向反转时效果也反转。总之,这些发现表明,高预测性能和测量声学因素中缺乏显著诊断组差异并不足以保证鲁棒性。我们认为,基于干预的鲁棒性测试应成为可信赖的临床语音模型的标准。
英文摘要
Speech-based Alzheimer's disease (AD) assessments increasingly rely on pretrained self-supervised learning (SSL) models that learn acoustic representations directly from raw audio, exposing the model to recording factors. We ask whether such factors are merely encoded in SSL representations or can systematically alter predictions. Using ADReSSo and three large SSL backbones, we apply controlled noise and reverberation interventions to participant-speech-only, non-speech, and full-recording audio. We combine layer-wise linear decoding, input- and representation-space interventions, and geometric alignment analysis to distinguish acoustic decodability from influence on AD prediction. Our results show that controlled acoustic interventions alter AD predictions across all three SSL backbones. Noise, despite showing no significant diagnostic-group difference in the original data, produces the strongest intervention effects. Importantly, these effects are systematically structured relative to the classifier's decision direction, replicate on the held-out test set and reverse when the representation-space intervention direction is reversed. Together, these findings show that high predictive performance and the absence of a significant diagnostic-group difference in a measured acoustic factor are not sufficient for robustness. We argue that intervention-based robustness tests should become standard for trustworthy clinical speech models.