使用基于语音的多模态大语言模型实现可泛化的认知障碍检测
Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models
另 1 家 · 查看机构详情
- Faculty of Digital Innovation, Arts \& Sciences, Saskatchewan Polytechnic, Regina SK S4S 5X1, Canada
- School of Basic Medical Sciences, Hebei University, Baoding 071000, China
- Department of Civil \& Environmental Engineering
- School of Mining \& Petroleum Engineering, University of Alberta, Edmonton AB T6G 2H5, Canada
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究针对认知障碍检测问题,提出基于开源大语言模型的多模态检测框架,整合语音音频与转录文本,提取特定模态嵌入后连接成特征向量用于分类,在相关基准数据集上评估,该框架准确率达92.4%,优于单模态基线且泛化能力强。
中文摘要 AI 辅助
认知障碍(CI)是一个日益严重的公共卫生问题。早期准确诊断对及时干预和改善患者预后至关重要。基于语音的CI检测是一种有前景的非侵入性方法,因为语音信号编码了与认知衰退相关的语言和声学标记。大语言模型(LLMs)的进展通过实现更具表现力的表征学习和跨不同说话者、记录设备及临床环境的更好泛化,增强了基于语音评估的潜力。此外,联合建模语言和声学特征的多模态学习能更全面地表征与CI相关的认知和行为变化,实现更可靠的检测。在这项工作中,我们提出了一个基于开源LLMs的多模态CI检测框架,该框架整合语音音频和相应转录文本,同时保护患者隐私。直接从语音信号中提取声学嵌入,从自动转录的语音中生成文本嵌入。然后将这些特定模态的嵌入连接起来创建组合特征向量,用于下游分类,无需访问原始或敏感患者数据。该方法在ADReSS20和ADReSSo21基准数据集上进行评估。实验结果表明,所提出的多模态框架实现了92.4%的CI分类准确率,始终优于单模态基线。我们的工作为CI识别建立了新的最先进水平,所提方法展示了卓越的跨数据集泛化能力。这一进展凸显了基于LLM的多模态框架融合语言和声学数据以实现强大、可扩展和非侵入性筛查的能力。
英文摘要
Cognitive impairment (CI) is a growing public health concern. Early and accurate diagnosis is critical for enabling timely intervention and improving patient outcomes. Speech-based CI detection has emerged as a promising non-invasive approach, as speech signals encode both linguistic and acoustic markers associated with cognitive decline. Recent advances in large language models (LLMs) further strengthen the potential of speech-based assessment by enabling more expressive representation learning and improved generalization across diverse speakers, recording devices, and clinical environments. Moreover, multimodal learning by jointly modeling linguistic and acoustic features allows for a more comprehensive characterization of cognitive and behavioral changes related to CI, leading to more reliable detection. In this work, we propose a multimodal CI detection framework based on open-source LLMs that integrates speech audio and corresponding transcripts while preserving patient privacy. Acoustic embeddings are extracted directly from speech signals, while textual embeddings are generated from automatically transcribed speech. These modality-specific embeddings are then concatenated to create a combined feature vector and used for downstream classification, without requiring access to raw or sensitive patient data. The proposed approach is evaluated on the ADReSS20 and ADReSSo21 benchmark datasets. Experimental results show that the proposed multimodal framework achieves an CI classification accuracy of 92.4% and consistently outperforms single-modality baselines. Our work establishes a new state-of-the-art for CI identification, with the proposed method demonstrating superior cross-dataset generalization. This advance highlights the power of an LLM-based multimodal framework that fuses linguistic and acoustic data to enable robust, scalable, and non-invasive screening.