发表机构
King's College London, University of London; Oxford Immune Algorithmics(伦敦大学国王学院; 牛津免疫算法公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究结合疾病特异性生物标志物图、ODE、CNN、LSTM及GPT-4o-mini RAG构建临床决策流程,在103种疾病的鉴别诊断中取得较高准确率,发现证据基础与诊断正确性存在解耦。
AI 中文摘要
诊断错误带来了沉重负担,而无约束的大型语言模型(LLM)仍易产生幻觉,且对定量实验室动态的整合能力较弱。我们开发了一种决策支持流程,结合疾病特异性生物标志物相关图、常微分方程(ODE)、深度序列分类以及检索增强生成(RAG)。针对全血细胞计数(FBC)库中的103种疾病类别,生物标志物网络被用作耦合矩阵,为每种疾病生成30条轨迹(共3090条)。一维卷积神经网络(CNN)和长短期记忆(LSTM)网络对疾病轨迹及6种动态簇进行分类。受限的GPT-4o-mini RAG层使用包含19种模式的BMJ最佳实践/NICE语料库生成鉴别诊断,并对其诊断适用性、证据基础及临床合理性进行评估。在5次随机种子运行中,CNN的疾病水平准确率为0.940±0.006(95%置信区间0.933--0.948),LSTM的准确率为0.852±0.019(95%置信区间0.828--0.875);CNN的优势为8.87个百分点(95%置信区间6.47--11.27;t(4)=10.26,p=5.1×10^-4;Hedges' g=3.67)。在100个抽样RAG案例中,96个成功解析;97.9%的案例引用了证据,71.9%的案例提及了真实诊断,综合评分为3.82/5,严格通过率为47.9%。核心发现是证据基础与诊断正确性之间存在解耦:分类器正确与分类器错误的输出在诊断适用性上存在差异,但在证据基础上无差异。事后分析证实,诊断评分存在1.02分的差异(曼-惠特尼检验p=0.0024;Hedges' g=0.72),而证据基础仅相差-0.02分(p=0.839;g=-0.04)。
英文摘要
Diagnostic error carries a burden, while unconstrained large language models (LLMs) remain vulnerable to hallucination and weak integration of quantitative laboratory dynamics. We developed a decision-support pipeline combining disease-specific biomarker correlation graphs, ordinary differential equations (ODEs), deep sequence classification, and retrieval-augmented generation (RAG). For 103 disease classes from a full blood count (FBC) repository, biomarker networks were used as coupling matrices to generate 30 trajectories per disease (3,090 total). A one-dimensional convolutional neural network (CNN) and long short-term memory (LSTM) network classified disease trajectories and six dynamical clusters. A constrained GPT-4o-mini RAG layer used a 19-pattern BMJ Best Practice/NICE corpus to generate differential diagnoses evaluated for diagnostic suitability, evidential grounding, and clinical plausibility. Across five random-seed runs, disease-level accuracy was $0.940 \pm 0.006$ for the CNN (95\% CI 0.933--0.948) and $0.852 \pm 0.019$ for the LSTM (95\% CI 0.828--0.875); the CNN advantage was 8.87 percentage points (95\% CI 6.47--11.27; $t(4)=10.26$, $p=5.1\times10^{-4}$; Hedges' $g=3.67$). Among 100 sampled RAG cases, 96 parsed successfully; evidence was cited in 97.9\%, the true diagnosis was mentioned in 71.9\%, and the composite score was 3.82/5 with a 47.9\% strict pass rate. The central finding was a decoupling between grounding and diagnostic correctness: classifier-correct versus classifier-wrong outputs differed in diagnostic suitability but not evidential grounding. Post-hoc analysis confirmed a 1.02-point diagnostic-score difference (Mann--Whitney $p=0.0024$; Hedges' $g=0.72$), whereas grounding differed by only $-0.02$ points ($p=0.839$; $g=-0.04$).
Comments14 pages