AD-CARE:一种基于指南、模态无关的LLM代理,用于多队列评估、公平性分析和读者研究的阿尔茨海默病诊断
AD-CARE: A Guideline-grounded, Modality-agnostic LLM Agent for Real-world Alzheimer's Disease Diagnosis with Multi-cohort Assessment, Fairness Analysis, and Reader Study
- The Hong Kong Polytechnic University(香港理工大学)
- Xuanwu Hospital, Capital Medical University(首都医科大学宣武医院)
- Amazon(亚马逊)
- Zhejiang University(浙江大学)
- Sany AI(三一人工智能)
- Yantai Yuhuangding Hospital, Qingdao University(青岛大学烟台毓璜顶医院)
- The University of Hong Kong(香港大学)
- The Third Affiliated Hospital of Sun Yat-sen University(中山大学附属第三医院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
AD-CARE通过整合临床指南和多模态数据,实现了多队列中84.9%的诊断准确率,显著提升诊断效率和公平性,适用于阿尔茨海默病的临床决策支持。
AI中文摘要:
阿尔茨海默病(AD)是随着人口老龄化而增长的全球健康挑战,及时准确的诊断对于减少个人和社会负担至关重要。然而,现实中的AD评估受到不完整、异质的多模态数据和不同地点及患者人口统计学差异的阻碍。尽管大型语言模型(LLMs)在生物医学领域显示出潜力,但其在AD中的应用大多局限于回答狭窄的、疾病特定的问题,而不是生成支持临床决策的综合性诊断报告。本文通过引入AD-CARE,一种模态无关的代理,扩展了LLM在临床决策支持中的能力,该代理能从不完整、异质的输入中进行指南导向的诊断评估,而无需填补缺失的模态。通过动态协调专门的诊断工具并嵌入临床指南到LLM驱动的推理中,AD-CARE生成透明的报告式输出,与现实中的临床工作流程一致。在包含10,303例的六个队列中,AD-CARE实现了84.9%的诊断准确率,相较于基线方法,诊断准确率提高了4.2%-13.7%。尽管队列层面存在差异,数据集特定的准确率仍保持稳健(80.4%-98.8%),并且代理始终优于所有基线方法。AD-CARE减少了种族和年龄亚组之间的性能差异,将四个指标的平均分散度减少了21%-68%和28%-51%。在受控的读者研究中,该代理提高了神经科医生和放射科医生的准确性,提高了6%-11%,并且将决策时间减少了一半以上。该框架相对于八个基础LLM,获得了2.29%-10.66%的绝对提升,并收敛了它们的性能。这些结果表明,AD-CARE是一种可扩展、实际可部署的框架,可以集成到常规的临床工作流程中,用于阿尔茨海默病的多模态决策支持。
英文摘要:
Alzheimer's disease (AD) is a growing global health challenge as populations age, and timely, accurate diagnosis is essential to reduce individual and societal burden. However, real-world AD assessment is hampered by incomplete, heterogeneous multimodal data and variability across sites and patient demographics. Although large language models (LLMs) have shown promise in biomedicine, their use in AD has largely been confined to answering narrow, disease-specific questions rather than generating comprehensive diagnostic reports that support clinical decision-making. Here we expand LLM capabilities for clinical decision support by introducing AD-CARE, a modality-agnostic agent that performs guideline-grounded diagnostic assessment from incomplete, heterogeneous inputs without imputing missing modalities. By dynamically orchestrating specialized diagnostic tools and embedding clinical guidelines into LLM-driven reasoning, AD-CARE generates transparent, report-style outputs aligned with real-world clinical workflows. Across six cohorts comprising 10,303 cases, AD-CARE achieved 84.9% diagnostic accuracy, delivering 4.2%-13.7% relative improvements over baseline methods. Despite cohort-level differences, dataset-specific accuracies remain robust (80.4%-98.8%), and the agent consistently outperforms all baselines. AD-CARE reduced performance disparities across racial and age subgroups, decreasing the average dispersion of four metrics by 21%-68% and 28%-51%, respectively. In a controlled reader study, the agent improved neurologist and radiologist accuracy by 6%-11% and more than halved decision time. The framework yielded 2.29%-10.66% absolute gains over eight backbone LLMs and converges their performance. These results show that AD-CARE is a scalable, practically deployable framework that can be integrated into routine clinical workflows for multimodal decision support in AD.