基于自发语音的可解释且可泛化的LLM认知衰退检测
Explainable and Generalisable LLM-based Cognitive Decline Detection with Spontaneous Speech
- Tsinghua University(清华大学)
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- Peking University Sixth Hospital(北京大学第六医院)
- Peking University(北京大学)
- OPPO Research Institute(OPPO研究院)
- Peking University Third Hospital(北京大学第三医院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出双语语音大语言模型框架,直接处理原始语音进行认知状态分类并生成可解释的自然语言说明,在多个数据集上取得最优性能,且具备跨任务泛化能力。
AI中文摘要:
阿尔茨海默病(AD)及其可能的前驱状态轻度认知障碍(MCI),会通过细微的语言和声学改变早期显现。然而,传统诊断方法往往资源密集且缺乏大规模筛查的可扩展性。为应对这些挑战,我们提出了一种新颖的双语语音大语言模型框架,用于自动化、可解释的认知筛查。与依赖易出错的自动语音识别(ASR)的传统流程不同,我们的系统直接处理原始语音,学习联合的声学-语义表征,从而保留转录中常丢失的关键韵律线索。利用我们新收集的PUTH-AD数据集及多个开源语料库,我们实现了多任务学习目标,该目标同时执行认知状态分类并生成临床医生可理解的自然语言解释。在六个数据集/任务条件下,与三个代表性基线相比,我们的系统取得了最高的平均准确率和AUROC。该系统展示了向留出的PUTH-AD任务子集的跨任务迁移能力,在完全未见过的认知任务上无需任务特定微调即可保持分类准确性。此外,临床医生评估证实,生成的解释既具有临床相关性,又与潜在的语音证据基本一致,支持其在临床解读中的潜在实用性。本研究为基于语音的认知筛查提供了一个可扩展、客观且可解释的框架,将认知状态分类与临床医生可评估和验证的自然语言解释相结合,弥合了先进AI与临床实用性之间的鸿沟。
英文摘要:
Alzheimer's disease (AD) and mild cognitive impairment (MCI), which may precede AD, manifest early through subtle linguistic and acoustic alterations. Traditional diagnostics, however, are often resource-intensive and lack scalability for mass screening. To address these challenges, we introduce a novel bilingual speech large language model framework for automated, explainable cognitive screening. Unlike conventional pipelines that rely on error-prone automatic speech recognition, our system directly processes raw speech to learn joint acoustic-semantic representations, preserving critical prosodic cues often lost in transcription. Utilising our newly collected PUTH-AD dataset alongside multiple open-source corpora, we implemented a multi-task learning objective that simultaneously performs cognitive status classification and generates clinician-understandable natural language explanations. Our system achieved the highest average accuracy and AUROC across six dataset/task conditions, comparing three representative baselines. The system demonstrated cross-task transfer to held-out PUTH-AD task subsets, maintaining classification accuracy on an entirely unseen cognitive task without task-specific fine-tuning. Furthermore, clinician evaluation confirms that the generated explanations are both clinically relevant and largely consistent with the underlying speech evidence, supporting their potential utility in clinical interpretation. This study provides a scalable, objective, and explainable framework for speech-based cognitive screening, combining cognitive status classification with natural language explanations that clinicians can assess and verify, bridging the gap between advanced AI and clinical utility.