arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.08338cs.CVcs.AIcs.LG

检索增强的视觉语言模型用于多模态黑色素瘤诊断

Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis

  • Handong Global University(-handong全球大学)

机构由 AI 辅助整理,请以论文原文为准。

Jihyun Moon, Charmgil Hong

更新

AI总结:

针对黑色素瘤诊断中VLM缺乏临床特异性的问题,提出检索增强VLM框架,通过纳入相似病例提示实现无需微调的准确分类与错误纠正,显著优于传统基线。

AI中文摘要:

恶性黑色素瘤的准确早期诊断对于改善患者预后至关重要。尽管卷积神经网络(CNNs)在皮肤镜图像分析中展现出潜力,但它们常常忽略临床元数据,并且需要大量的预处理。视觉语言模型(VLMs)提供了一种多模态替代方案,但在通用领域数据上训练时难以捕捉临床特异性。为了解决这一问题,我们提出了一种检索增强的VLM框架,该框架将语义相似的患者病例纳入诊断提示中。我们的方法无需微调即可实现有依据的预测,并显著提高了分类准确性和错误纠正能力,优于传统基线。这些结果表明,检索增强提示为临床决策支持提供了一种稳健的策略。

英文摘要:

Accurate and early diagnosis of malignant melanoma is critical for improving patient outcomes. While convolutional neural networks (CNNs) have shown promise in dermoscopic image analysis, they often neglect clinical metadata and require extensive preprocessing. Vision-language models (VLMs) offer a multimodal alternative but struggle to capture clinical specificity when trained on general-domain data. To address this, we propose a retrieval-augmented VLM framework that incorporates semantically similar patient cases into the diagnostic prompt. Our method enables informed predictions without fine-tuning and significantly improves classification accuracy and error correction over conventional baselines. These results demonstrate that retrieval-augmented prompting provides a robust strategy for clinical decision support.

补充信息

↑