arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2507.22902cs.HCcs.AIcs.CLcs.MA

迈向自主AI医生:真实世界场景下自主智能体AI与委员会认证临床医生的定量基准测试

Toward the Autonomous AI Doctor: Quantitative Benchmarking of an Autonomous Agentic AI Versus Board-Certified Clinicians in a Real World Setting

  • Doctronic Research Group(Doctronic研究组)
  • University of California San Francisco(加州大学旧金山分校)

机构由 AI 辅助整理,请以论文原文为准。

Hashim Hayat, Maksim Kudrautsau, Evgeniy Makarov, Vlad Melnichenko, Tim Tsykunou, Piotr Varaksin, Matt Pavelle, Adam Z. Oskowitz

更新

AI总结:

针对全球医疗人力短缺问题,研究验证多智能体AI系统Doctronic在虚拟急诊中的表现,发现其诊断、治疗方案一致性高,部分表现优于人类医生,为缓解医疗人力不足提供潜在方案。

AI中文摘要:

研究背景:预计到2030年,全球医疗从业者缺口将达1100万,且行政事务占用了50%的临床时间。人工智能(AI)有望助力缓解这些问题,但目前尚无基于大语言模型(LLM)的端到端自主AI系统在真实临床实践中接受过严格评估。本研究评估了基于多智能体LLM的AI框架能否在虚拟急诊场景中作为AI医生自主运作。研究方法:我们回顾性对比了多智能体AI系统Doctronic与委员会认证临床医生在500例连续急诊远程医疗接诊中的表现,主要终点包括诊断一致性、治疗方案一致性及安全性指标,通过设盲的基于LLM的裁定和人类专家评审进行评估。研究结果:Doctronic与临床医生的首要诊断在81%的病例中匹配,治疗方案在99.2%的病例中一致,未出现临床幻觉(即诊断或治疗无临床依据的情况)。在对不一致病例的专家评审中,AI表现更优的占36.1%,人类表现更优的占9.3%,其余病例诊断水平相当。研究结论:在本次针对自主AI医生的首次大规模验证中,我们证实其与人类临床医生的诊断及治疗方案具有高度一致性,AI表现与执业临床医生相当,部分情况下甚至更优。这些发现表明,多智能体AI系统可实现与人类医护人员相当的临床决策能力,为解决医疗人力短缺问题提供了潜在方案。

英文摘要:

Background: Globally we face a projected shortage of 11 million healthcare practitioners by 2030, and administrative burden consumes 50% of clinical time. Artificial intelligence (AI) has the potential to help alleviate these problems. However, no end-to-end autonomous large language model (LLM)-based AI system has been rigorously evaluated in real-world clinical practice. In this study, we evaluated whether a multi-agent LLM-based AI framework can function autonomously as an AI doctor in a virtual urgent care setting. Methods: We retrospectively compared the performance of the multi-agent AI system Doctronic and board-certified clinicians across 500 consecutive urgent-care telehealth encounters. The primary end points: diagnostic concordance, treatment plan consistency, and safety metrics, were assessed by blinded LLM-based adjudication and expert human review. Results: The top diagnosis of Doctronic and clinician matched in 81% of cases, and the treatment plan aligned in 99.2% of cases. No clinical hallucinations occurred (e.g., diagnosis or treatment not supported by clinical findings). In an expert review of discordant cases, AI performance was superior in 36.1%, and human performance was superior in 9.3%; the diagnoses were equivalent in the remaining cases. Conclusions: In this first large-scale validation of an autonomous AI doctor, we demonstrated strong diagnostic and treatment plan concordance with human clinicians, with AI performance matching and in some cases exceeding that of practicing clinicians. These findings indicate that multi-agent AI systems achieve comparable clinical decision-making to human providers and offer a potential solution to healthcare workforce shortages.

↑