arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

医学中的智能体人工智能:临床转化的架构、应用、评估及挑战

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

Zheng Tong, Yang Liu, Wanshu Fan, Jing Qin, Zhongbin Han, Haifan Gong, Congyu Liao, Xiaofeng Liu, Cong Wang

arXiv 2607.25489首次发表:更新:

发表机构

School of Software Engineering, Dalian University; Affiliated Zhongshan Hospital of Dalian University; The Chinese University of Hong Kong, Shenzhen; University of California, San Francisco; Yale University; The Hong Kong Polytechnic University(大连大学软件工程学院; 大连大学附属中山医院; 香港中文大学(深圳); 加州大学旧金山分校; 耶鲁大学; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨医学中智能体人工智能的架构、应用等,通过范围审查筛选相关研究,指出其范围未定论且评估与临床需求不符,证据基础有限,临床转化需更清晰定义、可重复评估等。

AI 中文摘要

大语言模型和多模态基础模型使医学人工智能系统超越孤立预测,能执行多步骤临床任务。但医学中智能体人工智能的范围尚无定论,当前评估实践与临床使用要求不符。我们进行了范围审查,通过五个电子来源进行系统证据映射,筛选1649条可导出记录,初步纳入557项符合预定义标准的独特研究。这些研究描述了使用外部工具的单智能体、由检索和外部知识支持的工作流程、多模态智能体以及应用于医学问答、图像解释等的多智能体系统。证据基础仍以公共基准、模拟设置等为主,对过程可靠性等评估不一致。临床转化依赖更清晰定义、可重复评估等。

英文摘要

Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and undertake multistep clinical tasks that require planning, tool use, memory, iterative correction, and coordination among specialized agents. However, the scope of agentic AI in medicine remains unsettled, and current evaluation practices are not yet aligned with the requirements of clinical use. We conducted a scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent collaboration. The included studies describe single agents that use external tools, workflows supported by retrieval and external knowledge, multimodal agents, and multi-agent systems applied to medical question answering, image interpretation, electronic health record analysis, drug safety, and clinical trial prediction. The evidence base remains dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation. Process reliability, evidence traceability, uncertainty, safety, workflow impact, and external validity are evaluated less consistently. Clinical translation will depend on clearer definitions, reproducible evaluation, auditable oversight, interoperable system design, and prospective validation in real-world clinical workflows.

CommentsReview article, 6 figures, 2 tables. 42 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑