发表机构
School of Health & Wellbeing, University of Glasgow(格拉斯哥大学健康与福祉学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于行为的融合模型,将LLMs与本体排名器结合,在不丢失证据的前提下提升罕见病诊断性能,在多个数据集和模型上实现了Recall@1的显著提升。
AI 中文摘要
本体排名器对罕见病诊断仍然有用,因为每个候选疾病都可以追溯到匹配的患者表型;大型语言模型(LLMs)可从同一患者描述生成鉴别诊断,但预测缺乏同样清晰的证据链。本文未探讨哪种系统应取代另一种,而是研究LLM能否在不放弃自身证据的前提下提升排名器性能。所提出的基于行为的融合模型会分析两个排名列表、它们的一致性以及每个候选疾病背后的本体支持,并学习针对单个病例如何依赖每个系统。对比前,研究人员移除了基准病例与本体注释源自同一出版物导致的已记录测试集泄漏路径。在8个开源LLMs上,融合方法在Phenopacket Store数据集上将Phenomizer的Recall@1提升7.86个百分点,在RAMEDIS数据集上提升20.18个百分点;当通过API与DeepSeek-V4-Flash配对时,仅在其他LLMs上训练的融合模型无需重新训练,就将Recall@1从0.1657提升至0.2176,增幅5.19个百分点。在90.8%的正确融合诊断中,疾病保留了可检查的候选级别本体证据。这些结果表明,LLMs可在不丢弃使其有用的结构化证据的前提下,强化成熟的诊断工具。
英文摘要
Ontology rankers remain useful for rare-disease diagnosis because each candidate can be traced to matched patient phenotypes. Large language models (LLMs) can generate differential diagnoses from the same patient description, but their predictions lack an equally clear evidence trail. Rather than asking which system should replace the other, we ask whether an LLM can improve the ranker without giving up its evidence. Our behavior-based fusion model examines the two ranked lists, their agreement, and the ontology support behind each candidate, and learns how much to rely on each system for the individual case. Before comparison, we remove a documented test-set leakage pathway caused by benchmark cases and ontology annotations being derived from the same publications. Across eight open LLMs, fusion improves Phenomizer Recall@1 by 7.86 percentage points on Phenopacket Store and 20.18 points on RAMEDIS. When paired with DeepSeek-V4-Flash through an API, a fusion model trained only on the other LLMs improves Recall@1 from 0.1657 to 0.2176, a 5.19-point gain, without retraining. For 90.8% of correct fused diagnoses, the disease retains candidate-level ontology evidence that can be inspected. These results show that LLMs can strengthen an established diagnostic tool without discarding the structured evidence that makes it useful.