AI 中文总结
本研究评估Claude Opus 4等5款大型语言模型对结直肠息肉的光学诊断性能,发现其区分息肉亚型的准确率接近专家共识,但灵敏度与特异度未达ESGE标准,需进一步研究方可临床应用。
AI 中文摘要
背景与研究目的:结直肠息肉的准确光学诊断可指导切除策略与随访监测,多模态大型语言模型(MLLMs)在基于图像的诊断中展现出潜力。本研究旨在评估MLLMs在结直肠息肉分类及组织学预测中的诊断准确性。方法:我们利用PRIME数据集开展回顾性诊断性能研究,该数据集包含经整理的白光与窄带成像(NBI)图像。针对Claude Opus 4、Google Gemini 2.5 Pro、GPT-o3、GPT-4o、GPT-5这5款MLLMs,对132例病例,参照专家回复,计算其在Paris分类、窄带成像结直肠内镜(NICE)分类及预测组织学方面的F1分数、正确百分比得分与准确率;采用Cochran Q检验与McNemar检验确定各MLLM预测值间的差异。结果:在肿瘤性与非肿瘤性息肉的分类中,所有MLLMs的F1分数均大于0.9;Gemini 2.5 Pro在浸润性与非浸润性息肉、低级别与高级别腺瘤分类中表现最佳,F1分数分别为0.560与0.492;采用Paris分类时,Claude Opus 4与GPT-5的正确百分比得分显著高于其他MLLMs,为41.7%。结论:Claude Opus 4与Gemini 2.5 Pro在区分息肉亚型时准确性最高,表现最接近专家共识;但其灵敏度与特异度未达到欧洲胃肠内镜学会(ESGE)标准,提示临床部署前需开展前瞻性多中心试验并设计人在回路的工作流程。
英文摘要
Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis. We aimed to evaluate the diagnostic accuracy of MLLMs in classifying colorectal polyps and predicting histology. Methods: We conducted a retrospective diagnostic performance study using the PRIME dataset, a curated set of white light and narrow-band imaging (NBI) images. We evaluated Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, and GPT-5. For Paris, Narrow-band Imaging Colorectal Endoscopic (NICE), and predicted histology, we calculated F1 scores, percent correct scores, and accuracy of each MLLM compared to expert responses for 132 cases. Cochran's Q and McNemar's Test were used to determine differences between predicted values of each MLLM. Results: The F1 scores among MLLMs were >0.9 for all models for neoplastic vs. non-neoplastic polyps. Gemini 2.5 Pro demonstrated the highest F1 scores for invasive vs. non-invasive polyps and low- vs. high-grade adenoma, at 0.560 and 0.492 respectively. Claude Opus 4 and GPT-5 had statistically significantly higher percent correct scores than other MLLMs at 41.7%, using Paris classification. Conclusions: Claude Opus 4 and Gemini 2.5 Pro showed the highest accuracy in differentiating polyp subtypes, performing closest to expert consensus. Sensitivity and specificity, however, did not meet ESGE standards, highlighting the need for prospective multicenter trials and the design of human-in-the-loop workflows before clinical deployment.
Comments22 pages, 1 figure, 5 tables