arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.04421cs.CLcs.AIcs.MA

多供应商多智能体大语言模型是否能改善临床诊断?

Do Mixed-Vendor Multi-Agent LLMs Improve Clinical Diagnosis?

Grace Chang Yuan, Xiaoman Zhang, Sung Eun Kim, Pranav Rajpurkar

更新

AI总结:

本文研究了多供应商多智能体系统在临床诊断中的表现,发现混合供应商配置在召回率和准确性上优于单一供应商配置,揭示了供应商多样性对诊断系统鲁棒性的重要性。

AI中文摘要:

多智能体大语言模型(LLM)系统已 emerge 为临床诊断的有前途的方法,通过代理之间的协作来细化医疗推理。然而,现有大多数框架依赖于单供应商团队(例如,同一模型家族的多个代理),这可能导致相关故障模式,强化共享偏见而非纠正它们。我们通过比较单LLM、单供应商和混合供应商多智能体对话(MAC)框架,研究了供应商多样性的影响。使用三个医生代理,分别实例化为o4-mini、Gemini-2.5-Pro和Claude-4.5-Sonnet,我们在RareBench和DiagnosisArena上评估性能。混合供应商配置始终优于单供应商配置,达到最先进的召回率和准确性。重叠分析揭示了其机制:混合供应商团队互补归纳偏见,揭示个体模型或同质团队集体遗漏的正确诊断。这些结果突显了供应商多样性作为稳健临床诊断系统关键设计原则的重要性。

英文摘要:

Multi-agent large language model (LLM) systems have emerged as a promising approach for clinical diagnosis, leveraging collaboration among agents to refine medical reasoning. However, most existing frameworks rely on single-vendor teams (e.g., multiple agents from the same model family), which risk correlated failure modes that reinforce shared biases rather than correcting them. We investigate the impact of vendor diversity by comparing Single-LLM, Single-Vendor, and Mixed-Vendor Multi-Agent Conversation (MAC) frameworks. Using three doctor agents instantiated with o4-mini, Gemini-2.5-Pro, and Claude-4.5-Sonnet, we evaluate performance on RareBench and DiagnosisArena. Mixed-vendor configurations consistently outperform single-vendor counterparts, achieving state-of-the-art recall and accuracy. Overlap analysis reveals the underlying mechanism: mixed-vendor teams pool complementary inductive biases, surfacing correct diagnoses that individual models or homogeneous teams collectively miss. These results highlight vendor diversity as a key design principle for robust clinical diagnostic systems.

补充信息

↑