发表机构
Center for Language and Speech Processing (CLSP), Johns Hopkins University; Research Center for Information Technology Innovation, Academia Sinica; University of Michigan, Ann Arbor; Carnegie Mellon University(约翰斯·霍普金斯大学语言与语音处理中心; 中央研究院资讯科技创新研究中心; 密歇根大学安娜堡分校; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究重新审视语音预处理与数据集整理对阿尔茨海默病检测的影响,发现经语音增强的数据集虽提升域内性能但降低跨域鲁棒性,更清晰的语音数据集未必更可靠。
AI 中文摘要
基于语音的阿尔茨海默病(AD)检测越来越依赖皮特语料库(Pitt Corpus)经语音增强和整理后的版本,其中语音增强、样本选择和人口统计学平衡通常被视为有益的预处理步骤。然而,这些转换是提升了现实世界中的AD检测效果,还是会影响模型的泛化能力和预测行为,目前尚不清楚。在本研究中,我们重新审视了语音预处理和数据集整理在广泛使用的基于语音的AD检测基准中的作用。我们评估了不同数据集的语音质量、多种深度学习模型在匹配和不匹配增强设置下的跨数据集泛化能力,以及几种近期大型音频语言模型(LALMs)的行为。实验结果表明,在多个监督语音模型中,经语音增强的数据集通常能提升域内性能,但会降低跨域评估的鲁棒性;训练数据与测试数据间的匹配增强可缓解但无法消除这种性能下降。大型音频语言模型表现出类似的敏感性:与未处理数据相比,增强数据集会引发更严重的类别不平衡和预测偏移。这些结果表明,语音预处理和数据集整理会显著影响下游AD检测的行为,意味着“更清晰”的语音数据集不一定对现实世界的AD检测更可靠。
英文摘要
Speech-based Alzheimer's disease (AD) detection increasingly relies on speech-enhanced and curated versions of the Pitt Corpus, where speech enhancement, sample selection, and demographic balancing are often treated as beneficial preprocessing steps. However, whether these transformations improve real-world AD detection or instead affect model generalization and prediction behavior remains unclear. In this work, we revisit the role of speech preprocessing and dataset curation across widely used benchmarks for speech-based AD detection. We evaluate the speech quality of different datasets, the cross-dataset generalization of multiple deep learning models under matched and mismatched enhancement settings, and the behavior of several recent large audio-language models (LALMs). Experimental results show that across multiple supervised speech models, speech-enhanced datasets often improve in-domain performance while reducing robustness in cross-domain evaluation. Matched enhancement between training and test data alleviates, but does not eliminate, this degradation. LALMs show a similar sensitivity: enhanced datasets induce stronger class imbalance and prediction shifts than unprocessed data. These results suggest that speech preprocessing and dataset curation can substantially influence downstream AD detection behavior, indicating that ``cleaner'' speech datasets are not necessarily more reliable for real-world AD detection.