arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35549cs.AI

RareDx:面向罕见病诊断的受控知识整合与图基策略优化

RareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis

Bo Zhang, Yuchen Wang, Dongbai Li, Matthew Yu Heng Wong, Qingkai Zeng, Lijun Wang, Tien-Yin Wong, Peng Cui, Tianyu Liu

首次发表
浏览论文内容

中文总结 AI 辅助

RareDx通过受控知识整合与知识图谱基策略优化,将紧凑模型提升为罕见病长尾诊断的竞争力排序器,在八个基准上超越GPT-5.5。

中文摘要 AI 辅助

罕见病诊断是一个长尾推理问题:表型信息不完整,单个疾病记录稀疏,相关证据分布在本体论、基因注释和生物医学文本中。因此,语言模型倾向于常见疾病,遗漏罕见候选,或产生看似合理但无效的名称。我们提出RareDx,它将受控证据使用与知识图谱基策略优化相结合。RareDx-Harness将异构记录规范化为一个排序诊断任务,并在共享知识层上比较直接推理、静态检索、自适应工具和结构化表型-基因-疾病推理。训练流程将Top-10后训练与我们的知识图谱基策略优化方法RareDx-KGPO相结合。其奖励将预测投影到规范疾病图谱中,并整合了策划的分级相关性、本体邻近性、生物医学相似性和表型一致性。词汇和输出预算约束防止密集的部分信用奖励虚构或过长的鉴别诊断。在八个基准测试中,以Qwen3.5-9B为核心的完整RareDx系统在归档协议下达到38.34的宏Hit@10,比GPT-5.5高出1.60个百分点;一项不相交的验证选择审计在保留病例上保留了比Direct高6.80个百分点的路由增益。27B系统在Hit@1/5/10上达到23.53/36.56/40.76。受控消融实验表明,检索并非普遍有用,受控路由是增益的核心。这些结果表明,结构化医学知识可以将紧凑模型转变为在临床实践中异构长尾设置下具有竞争力的诊断排序器。

英文摘要

Rare-disease diagnosis is a long-tail reasoning problem: phenotypes are incomplete, individual disorders are sparsely documented, and relevant evidence is distributed across ontologies, gene annotations, and biomedical text. Language models consequently favor common conditions, miss rare candidates, or produce plausible but invalid names. We introduce RareDx, which couples controlled evidence use with knowledge-graph-grounded policy optimization. RareDx-Harness normalizes heterogeneous records into one ranked-diagnosis task and compares direct inference, static retrieval, adaptive tools, and structured phenotype-gene-disease reasoning over a shared knowledge layer. The training pipeline combines Top-10 post-training with RareDx-KGPO, our knowledge-graph-grounded policy optimization method. Its reward projects predictions into a canonical disease graph and integrates curated graded relevance, ontology proximity, biomedical similarity, and phenotype consistency. Vocabulary and output-budget constraints prevent dense partial credit from rewarding fabricated or overlong differentials. Across eight benchmarks, the complete RareDx system centered on Qwen3.5-9B reaches 38.34 macro Hit@10, 1.60 points above GPT-5.5 under the archived protocol; a disjoint validation-selection audit retains a 6.80-point routing gain over Direct on held-out cases. The 27B system reaches 23.53/36.56/40.76 at Hit@1/5/10. Controlled ablations show that retrieval is not uniformly helpful and that controlled routing is central to the gain. These results indicate that structured medical knowledge can turn a compact model into a competitive diagnostic ranker across heterogeneous long-tail settings in clinical practice.

发表机构

  • UIUC(伊利诺伊大学厄巴纳-香槟分校)
  • Tsinghua University(清华大学)
  • University of Cambridge(剑桥大学)
  • Nankai University(南开大学)
  • Zhejiang University(浙江大学)
  • Yale University(耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑