arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24190cs.CV

基准测试:现成多模态AI模型与皮肤科医生在患者拍摄皮肤图像上的比较

Benchmarking Off-the-Shelf Multimodal AI Models Against Dermatologists on Patient-Captured Skin Images

Rian Dolphin, Laura Knowles

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估三个低价多模态AI模型在患者皮肤图像诊断中的表现,发现其与皮肤科医生一致性相当,但置信度校准差,元数据影响因模型而异,且成本不预示性能。

中文摘要 AI 辅助

人工智能(AI)近年来发展迅速。最初,大型语言模型的突破引起了广泛关注。然而,最近几代前沿AI模型已将多模态能力作为一等公民,其中视觉能力是核心。在本文中,我们评估了三个最近发布的模型在从患者提交的图像中诊断皮肤病状况的任务上的表现。所选模型在定价上属于低到中档,因此代表了当前AI能力的最低水平,而非最高水平。我们将AI性能与三位认证皮肤科医生组成的专家组进行比较,他们对每张图像进行评分,并提出了四个有趣的发现。首先,根据不同的指标,测试的AI模型在临床医生间一致性方面与人类相当或略逊于人类。其次,我们发现要求AI模型提供置信度评级会产生校准不良的答案,这意味着在临床环境中不应依赖置信度阈值的使用。第三,提供额外患者元数据的效果高度依赖模型,三个模型中的一个在每项指标上均出现性能下降。最后,模型成本并不能预测性能。我们测试的最佳性能模型平均每个病例花费0.0045美元。

英文摘要

Artificial intelligence (AI) has advanced at a rapid pace in recent years. Initially, breakthroughs in large language models caught widespread attention. However, recent generations of frontier AI models have adopted multimodal capabilities as a first class citizen, with vision capabilities being central to that. In this paper, we evaluate three recently released models on the task of diagnosing dermatological conditions from patient-submitted images. The models chosen are at the low to mid tier in terms of pricing and thus represent a floor on current AI capabilities, not a ceiling. We evaluate AI performance relative to a panel of three certified dermatologists, who grade each image, and we present four interesting findings. Firstly, depending on the metric, the tested AI models are either on par or slightly trail humans in terms of inter-clinician agreement. Secondly, we find that asking AI models for a confidence rating produces poorly calibrated answers, meaning use of confidence thresholds should not be relied upon in a clinical setting. Thirdly, the effect of providing additional patient metadata is strongly model-specific, with one of the three models degrading on every metric considered. Finally, model cost is not predictive of performance. The best-performing model we tested costs on average $0.0045 per case.

发表机构

  • University of Limerick(利默里克大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑