arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种智能体AI框架克服了大型语言模型在眼底照相青光眼检测中的基本局限

An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography

Jalil Jalili, Hossein Taghizad, Anuwat Jiravarnsirikul, Christopher Bowd, Akram Belghith, Raheleh Kafieh, Christopher A. Girkin, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher

arXiv 2608.07651首次发表:更新:

AI 中文总结

该研究开发的智能体AI框架整合LLM与专用深度学习工具,提升了眼底照相青光眼检测的准确率、一致性,纠正了仅用LLM的缺陷,具有通用性,或推动医学AI向多智能体系统转变。

AI 中文摘要

大型语言模型(LLM)在医学影像解读方面展现出应用前景,但存在幻觉、准确率有限及运行间不一致性等问题。我们开发并验证了一种将LLM与专用深度学习工具结合的智能体AI框架,用于眼底照相的青光眼检测。该工作流程包含三个步骤:(1)LLM初始评估;(2)函数调用以调用专用工具,分别用于图像质量检测(QAModel、FundaQ-8)、青光眼分类(SwinV2-Tiny)以及视盘/视杯分割(SegFormer-B0);(3)LLM反思,整合初始评估与工具输出。我们在两个公共数据集(ORIGA,n=100;RIM-ONE-v3,n=100)上,对未裁剪和裁剪视野的图像,评估了两种LLM(Gemini 2.5 Flash、GPT-5.4 mini),所有图像均由一名经盲法处理的、接受过专科培训的青光眼专家独立分级。在所有条件下,智能体工作流程将分类准确率提升了16至47个百分点,达到与专家准确率相差6个百分点以内;在RIM-ONE-v3数据集上,最佳配置达到了专家的88%准确率。仅使用LLM的方法存在两类缺陷:GPT-5.4 mini呈现阳性偏差(灵敏度95-100%,特异性0-5%),而Gemini 2.5 Flash在不同运行间表现出随机波动;智能体工作流程纠正了这两类问题。视杯视盘比误差降低了15-50%(平均绝对误差从0.156-0.228降至0.104-0.132),与专家分级的相关性从弱相关(r=0.12-0.39)提升至中强相关(r=0.59-0.84),运行间一致性从接近随机(Cohen’s kappa低至-0.01)提升至接近完美(Cohen’s kappa高达0.96)。将LLM与专用工具结合解决了仅用LLM方法的关键局限,包括过度诊断和运行间变异性。该提升对两种LLM均成立,表明其在不同主干模型间具有通用性,或标志着医学AI从单一模型向协调式多智能体系统的转变。

英文摘要

Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framework integrating LLMs with specialized deep learning tools for glaucoma detection from fundus photography. The workflow had three steps: (1) LLM initial assessment; (2) function calling to invoke specialized tools for image quality (QAModel, FundaQ-8), glaucoma classification (SwinV2-Tiny), and optic disc/cup segmentation (SegFormer-B0); and (3) LLM reflection integrating the initial impression with tool outputs. Two LLMs (Gemini 2.5 Flash, GPT-5.4 mini) were evaluated on two public datasets (ORIGA, n=100; RIM-ONE-v3, n=100) under uncropped and cropped fields of view; all images were independently graded by a masked fellowship-trained glaucoma specialist. The agentic workflow improved classification accuracy by 16 to 47 percentage points across all conditions, reaching within 6 points of the specialist; on RIM-ONE-v3 the best configurations matched the specialist accuracy of 88%. LLM-alone approaches failed in two ways: GPT-5.4 mini showed positive bias (sensitivity 95-100%, specificity 0-5%), while Gemini 2.5 Flash varied stochastically between runs; the agentic workflow corrected both. Cup-to-disc ratio error fell 15-50% (MAE 0.156-0.228 to 0.104-0.132), and correlation with specialist grading rose from weak (r=0.12-0.39) to moderate-strong (r=0.59-0.84). Run-to-run consistency rose from near-random (kappa as low as -0.01) to near-perfect (kappa up to 0.96). Integrating LLMs with specialized tools addressed key limitations of LLM-alone approaches, including over-diagnosis and run-to-run variability. Gains held for both LLMs, suggesting generalizability across backbones, and may signal a shift from monolithic models toward orchestrated multi-agent systems in medical AI.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑