RareLens:通过对齐不同的大语言模型推理实现端到端罕见病护理
RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning
浏览论文内容
中文总结 AI 辅助
针对罕见病诊断难题,RareLens系统利用不同大语言模型推理的差异,通过四个协同模块实现全病程临床决策支持,在真实数据集和外部研究中表现出色,证明对齐模型推理是高不确定性临床决策的有效策略。
中文摘要 AI 辅助
罕见病影响约3.5%至5.9%的人口,超70%患者被误诊,因早期症状不特异且相关专业知识稀缺分布不均。现有人工智能系统多处理孤立护理阶段,依赖下游检查结果并视模型差异为噪声。本文提出RareLens系统,利用模型差异支持罕见病全病程临床决策。四个协同模块进行初诊风险筛查、诊断、治疗规划和预后评估。在RareBench数据集上开发和评估,RareLens在各阶段均优于测试的前沿模型,在外部研究中,自主的RareLens和受其辅助的医生均显著优于未受辅助的医生。这表明对齐不同模型推理为高不确定性临床决策提供了可推广策略。
英文摘要
Rare diseases represent one of the most challenging settings for clinical decision-making, where heterogeneous presentations, sparse evidence and limited expertise create persistent uncertainty throughout the care pathway. Although artificial intelligence could help, existing systems largely address isolated tasks, particularly diagnosis, and usually rely on downstream investigations rather than information available at initial presentation. Here we show that clinical AI performance under uncertainty can be improved not by scaling a single model, but by exploiting the diversity of multiple imperfect reasoning systems. Across heterogeneous large language models, we identify divergent reasoning trajectories with complementary error patterns and develop RareLens, which learns to reconcile these perspectives into actionable decisions across four stages of rare disease care: risk screening, diagnosis, treatment planning and prognosis prediction. Built on RarelensBench, a real-world dataset of 157,525 cases spanning all 33 Orphanet categories and more than 7,000 conditions, RareLens outperformed every frontier model tested, including GPT-5, DeepSeek-R1, Claude-3.7-Sonnet and Gemini-2.5-Pro, across all stages. It achieved an area under the curve of 0.917 for screening and top-1 accuracies of 65.5% and 89.8% for diagnosis and treatment. In an external evaluation involving 1,287 cases and 23 physicians, autonomous RareLens and physicians assisted by RareLens both outperformed unaided physicians, while demonstrating that effective human-AI collaboration requires more than simply providing model outputs. These findings establish divergent model reasoning as an exploitable source of information and suggest a general strategy for building AI systems that operate reliably under high clinical uncertainty.