arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向复杂临床推理的稀疏多阶段智能体专家路由机制

Sparse Multi-Stage Expert-Agent Routing for Complex Clinical Reasoning

Sike Xiang, Shuang Chen, Qian sun, Jia Cheng, Yusi Wei, Amir Atapour-Abarghouei

arXiv 2608.21948首次发表:更新:

发表机构

Durham University(杜伦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Sparse Multi-Stage Expert-Agent Routing框架,通过自适应激活稀疏专家智能体实现高效临床推理,引入ClinFEScore评估协议,在200例真实MDT病例上达91.5%诊断准确率,激活专家数从17.0降至3.0。

AI 中文摘要

复杂临床推理要求模型在新证据出现时更新诊断假设,并在有限会诊资源下协调不同医学专科。现有基于大语言模型(LLM)的临床推理系统通常执行单次预测或依赖固定多智能体工作流,导致专家参与要么静态要么不必要地详尽。本文提出Sparse Multi-Stage Expert-Agent Routing,一种基于语言的临床推理框架,将诊断建模为分阶段路由过程。给定从多模态逐步获取的临床证据,该框架维护不断演变的病例状态,并自适应激活稀疏的医学专家智能体集合,各阶段由专家专属记忆支持。为评估自由文本诊断结论而非表面相似性,本文进一步引入ClinFEScore,一种针对临床推理输出的事实感知语义评估协议。在从MAC和AgentClinic-NEJM重构的多阶段病例上,本文框架将激活专家的平均数量从17.0降至3.0,同时保持较强的事实级诊断质量。在200个真实世界医院多学科诊疗(MDT)病例上,ClinFEScore与临床医生判断具有强相关性(Spearman's ρ=0.81;Pearson's r=0.87),而本文方法实现了91.5%经临床医生验证的诊断准确率,每个病例仅需约5次专家智能体/LLM调用。这些结果表明,稀疏分阶段协调是一种高效且符合临床需求的基于LLM的临床推理方法。

英文摘要

Complex clinical reasoning requires models to update diagnostic hypotheses as new evidence emerges and to coordinate different medical specialities under limited consultation resources. Existing LLM-based clinical reasoning systems typically perform single-pass prediction or rely on fixed multi-agent workflows, making expert participation either static or unnecessarily exhaustive. We propose Sparse Multi-Stage Expert-Agent Routing, a language-based clinical reasoning framework that models diagnosis as a stage-wise routing process. Given progressively available clinical evidence derived from multiple modalities, the framework maintains an evolving case state and adaptively activates a sparse set of medical expert agents, supported by expert-specific memory across stages. To evaluate free-text diagnostic conclusions beyond surface similarity, we further introduce ClinFEScore, a fact-aware semantic evaluation protocol for clinical reasoning outputs. On reconstructed multi-stage cases from MAC and AgentClinic-NEJM, our framework reduces the average number of activated experts from 17.0 to 3.0 whilst maintaining strong fact-level diagnostic quality. On 200 real-world hospital MDT cases, ClinFEScore correlates strongly with clinician judgements (Spearman's $ρ=0.81$; Pearson's $r=0.87$), whilst our method achieves 91.5\% clinician-verified diagnostic accuracy with approximately five expert-agent/LLM calls per case. These results support sparse stage-wise coordination as an efficient and clinically relevant approach to LLM-based clinical reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑