发表机构
Xiamen University; National Institute for Data Science in Health and Medicine, Xiamen University; Institute of Artificial Intelligence, Xiamen University; School of Medicine, Xiamen University; The First Affiliated Hospital of Xiamen University(厦门大学; 厦门大学健康医疗大数据国家研究院; 厦门大学人工智能研究院; 厦门大学医学院; 厦门大学附属第一医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对表型驱动罕见病诊断基准测试不透明问题,引入GraphRareBench基准测试,含病例与目标-混杂因素对。通过共享特征接口监督排名器及特定智能体测试,发现不同方法互补,为更透明的诊断系统评估提供基础。
AI 中文摘要
表型驱动的诊断基准测试通常报告参考疾病的排名,但很少揭示哪些合理替代方案排名高于它,或者工具使用模型在做出决策前检查了哪些证据。我们引入了GraphRareBench,这是一个可溯源的基准测试,包含2365个源自本体的病例和18093个目标-混杂因素对。每个病例包括一个粗化的HPO查询、固定的候选池、图定义的硬混杂因素和源链接的证据记录。在237个病例的基因-组件不相交测试拆分中,使用共享21特征接口的监督排名器的平均倒数排名范围为0.640至0.740,病例平均目标-混杂因素准确率范围为0.898至0.916。使用Agents-A1和DeepSeek-V4-Flash实例化的智能体的平均倒数排名分别为0.746和0.718。它们的配对平均倒数排名差异无统计学意义,而目标证据覆盖率相差0.561。这些结果表明,全池检索、硬混杂因素辨别和可观察证据访问捕捉了模型行为的互补方面。GraphRareBench为表型驱动诊断系统的更透明和证据感知评估提供了基础。代码和数据可在指定网址获取。
英文摘要
Phenotype-driven diagnostic benchmarks usually report the rank of the reference disease, but they rarely reveal which plausible alternatives are ranked above it or what evidence a tool-using model examines before making its decision. We introduce GraphRareBench, a provenance-preserving benchmark containing 2,365 ontology-derived cases and 18,093 target-confounder pairs. Each case includes a coarsened HPO query, a fixed candidate pool, graph-defined hard confounders, and source-linked evidence records. On the 237-case gene-component-disjoint test split, supervised rankers using a shared 21-feature interface achieved MRRs ranging from 0.640 to 0.740 and case-averaged target-over-confounder accuracies ranging from 0.898 to 0.916. Agents instantiated with Agents-A1 and DeepSeek-V4-Flash achieved MRRs of 0.746 and 0.718, respectively. Their paired MRR difference was not statistically significant, whereas their target-evidence coverage differed by 0.561. Together with the observation that 22.1% to 43.7% of selected Hit@10 successes still ranked at least one graph-defined hard confounder above the target, these results indicate that full-pool retrieval, hard-confounder discrimination, and observable evidence access capture complementary aspects of model behavior. GraphRareBench therefore provides a foundation for more transparent and evidence-aware evaluation of phenotype-driven diagnostic systems. Code and data are available at https://github.com/GUI0609/GraphRareBench.