AI 中文总结
该研究开发的智能体AI框架整合LLM与专用深度学习工具,提升了眼底照相青光眼检测的准确率、一致性,纠正了仅用LLM的缺陷,具有通用性,或推动医学AI向多智能体系统转变。
AI 中文摘要
大型语言模型(LLM)在医学影像解读方面展现出应用前景,但存在幻觉、准确率有限及运行间不一致性等问题。我们开发并验证了一种将LLM与专用深度学习工具结合的智能体AI框架,用于眼底照相的青光眼检测。该工作流程包含三个步骤:(1)LLM初始评估;(2)函数调用以调用专用工具,分别用于图像质量检测(QAModel、FundaQ-8)、青光眼分类(SwinV2-Tiny)以及视盘/视杯分割(SegFormer-B0);(3)LLM反思,整合初始评估与工具输出。我们在两个公共数据集(ORIGA,n=100;RIM-ONE-v3,n=100)上,对未裁剪和裁剪视野的图像,评估了两种LLM(Gemini 2.5 Flash、GPT-5.4 mini),所有图像均由一名经盲法处理的、接受过专科培训的青光眼专家独立分级。在所有条件下,智能体工作流程将分类准确率提升了16至47个百分点,达到与专家准确率相差6个百分点以内;在RIM-ONE-v3数据集上,最佳配置达到了专家的88%准确率。仅使用LLM的方法存在两类缺陷:GPT-5.4 mini呈现阳性偏差(灵敏度95-100%,特异性0-5%),而Gemini 2.5 Flash在不同运行间表现出随机波动;智能体工作流程纠正了这两类问题。视杯视盘比误差降低了15-50%(平均绝对误差从0.156-0.228降至0.104-0.132),与专家分级的相关性从弱相关(r=0.12-0.39)提升至中强相关(r=0.59-0.84),运行间一致性从接近随机(Cohen’s kappa低至-0.01)提升至接近完美(Cohen’s kappa高达0.96)。将LLM与专用工具结合解决了仅用LLM方法的关键局限,包括过度诊断和运行间变异性。该提升对两种LLM均成立,表明其在不同主干模型间具有通用性,或标志着医学AI从单一模型向协调式多智能体系统的转变。
英文摘要
Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framework integrating LLMs with specialized deep learning tools for glaucoma detection from fundus photography. The workflow had three steps: (1) LLM initial assessment; (2) function calling to invoke specialized tools for image quality (QAModel, FundaQ-8), glaucoma classification (SwinV2-Tiny), and optic disc/cup segmentation (SegFormer-B0); and (3) LLM reflection integrating the initial impression with tool outputs. Two LLMs (Gemini 2.5 Flash, GPT-5.4 mini) were evaluated on two public datasets (ORIGA, n=100; RIM-ONE-v3, n=100) under uncropped and cropped fields of view; all images were independently graded by a masked fellowship-trained glaucoma specialist. The agentic workflow improved classification accuracy by 16 to 47 percentage points across all conditions, reaching within 6 points of the specialist; on RIM-ONE-v3 the best configurations matched the specialist accuracy of 88%. LLM-alone approaches failed in two ways: GPT-5.4 mini showed positive bias (sensitivity 95-100%, specificity 0-5%), while Gemini 2.5 Flash varied stochastically between runs; the agentic workflow corrected both. Cup-to-disc ratio error fell 15-50% (MAE 0.156-0.228 to 0.104-0.132), and correlation with specialist grading rose from weak (r=0.12-0.39) to moderate-strong (r=0.59-0.84). Run-to-run consistency rose from near-random (kappa as low as -0.01) to near-perfect (kappa up to 0.96). Integrating LLMs with specialized tools addressed key limitations of LLM-alone approaches, including over-diagnosis and run-to-run variability. Gains held for both LLMs, suggesting generalizability across backbones, and may signal a shift from monolithic models toward orchestrated multi-agent systems in medical AI.