arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2511.14900cs.CVcs.AIcs.CL

Skin-R1: 基于视觉-语言模型的临床知识引导的皮肤病诊断

Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models

Zehao Liu, Weijieying Ren, Jipeng Zhang, Tianxiang Zhao, Jingxi Zhu, Xiaoting Li, Vasant G Honavar

更新

AI总结:

Skin-R1结合教材引导的临床推理监督与强化学习,提升皮肤病诊断的准确性和鲁棒性,通过构建基于教材的推理生成器和强化学习框架,改进模型在异构数据集和稀疏标注下的表现。

AI中文摘要:

视觉-语言模型(VLMs)最近在辅助皮肤病诊断的临床推理中显示出潜力。然而,其可信度和临床实用性受限于三个关键挑战:异构数据集具有不一致的诊断标签和概念注释、缺乏可靠的推理监督的诊断理由、以及从小规模密集标注数据集向大规模稀疏标注数据集迁移知识时的扩展性有限。为此,我们提出了Skin-R1,一种面向皮肤病的VLM,整合了基于教科书的临床推理监督与强化学习(RL),以提高诊断预测的准确性和鲁棒性。首先,我们构建了一个基于教科书的推理生成器,该生成器从权威的皮肤病知识中合成具有层次意识和差异诊断(DDx)的诊断轨迹。其次,这些轨迹用于监督微调(SFT),建立模型的临床基础推理。最后,我们引入了一个强化学习训练框架,将皮肤病的层次结构融入奖励设计中,使模型能够将基于基础的诊断推理泛化到大规模稀疏标注的数据集上。在多个皮肤病基准测试中,广泛实验表明Skin-R1在诊断准确性和鲁棒性方面均优于最先进的Med-VLM基线。消融研究进一步突显了在SFT阶段引入的基于基础的推理监督的重要性。

英文摘要:

Vision--language models (VLMs) have recently shown promise for assisting clinical reasoning in dermatological diagnosis. However, their trustworthiness and clinical utility remain limited by three key challenges: heterogeneous datasets with inconsistent diagnostic labels and concept annotations, the lack of grounded diagnostic rationales for reliable reasoning supervision, and limited scalability when transferring knowledge from small, densely annotated datasets to large collections with sparse labels. To address these challenges, we propose Skin-R1, a dermatology-oriented VLM that integrates textbook-grounded clinical reasoning supervision with reinforcement learning (RL) to improve the accuracy and robustness of diagnostic prediction. First, we construct a textbook-based reasoning generator that synthesizes hierarchy-aware and differential-diagnosis (DDx) diagnostic trajectories derived from authoritative dermatology knowledge. Second, these trajectories are used for supervised fine-tuning (SFT), establishing a clinically grounded reasoning foundation for the model. Finally, we introduce an RL training framework that incorporates the hierarchical structure of dermatological diseases into the reward design, enabling the model to generalize grounded diagnostic reasoning to large-scale datasets with sparse annotations. Extensive experiments across multiple dermatology benchmarks demonstrate that Skin-R1 consistently improves diagnostic accuracy and robustness compared to state-of-the-art Med-VLM baselines. Ablation studies further highlight the critical role of grounded reasoning supervision introduced during the SFT stage.

补充信息

↑