发表机构
Tribhuvan University; North Carolina A&T State University; Aalto University; Auburn University(特里布万大学; 北卡罗来纳农工州立大学; 阿尔托大学; 奥本大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究分解皮肤病学AI的泛化差距,发现疾病分布偏移比肤色影响更大,提出皮肤病预训练特征的可转移性更强,轻量适应仅需少量标记示例即可恢复性能,并发布评估协议与代码。
AI 中文摘要
皮肤病学人工智能(AI)模型主要在浅色皮肤、聚焦癌症的图像集合上训练,却被越来越多地提议部署在资源受限的环境中,这些环境中的患者与训练人群在两个混淆维度上存在差异:肤色和疾病分布。我们研究泛化性能差是否主要由肤色代表性不足或疾病分布偏移导致。我们评估了一个在HAM10000和ISIC 2019上微调的癌症训练基线模型(ResNet-50)、两个皮肤病学基础模型(DermLIP和MONET),以及一个通用视觉模型(DINOv3)作为冻结特征提取器。模型在肤色分层的疾病匹配数据集(多样化皮肤病图像集,DDI)和疾病偏移的肤色多样化数据集(皮肤病图像网络,SCIN)上进行评估。我们的结果表明,在评估设置中,疾病分布偏移比肤色的影响更大。当转移到不熟悉的临床条件时,癌症基线的平衡准确率从0.62降至0.21,而疾病内的肤色差距更小(0.10-0.18)且不一致。无标签表示分析表明,这种失败反映了表示限制,而非仅缺少输出标签:癌症专用特征对不熟悉条件的聚类效果差(kNN纯度提升较随机水平为+0.06),而皮肤病预训练特征保留了更强的可转移结构(+0.23)。最后,我们表明表示质量可预测轻量适应下的可恢复性能。从皮肤病学基础模型出发,每个临床类别约10个标记示例即可恢复大部分可达到的性能。我们发布评估协议和代码,以支持皮肤病学AI泛化的可复现审计。
英文摘要
Dermatology artificial intelligence (AI) models are predominantly trained on light-skinned, cancer-focused image collections, yet they are increasingly proposed for deployment in resource-constrained settings where patients differ from training populations along two confounded axes: skin tone and disease distribution. We investigate whether poor generalization is primarily caused by skin-tone underrepresentation or disease-distribution shift. We evaluate a cancer-trained baseline (ResNet-50 fine-tuned on HAM10000 and ISIC 2019), two dermatology foundation models (DermLIP and MONET), and a general-purpose vision model (DINOv3) as frozen feature extractors. Models are evaluated on a tone-stratified disease-matched dataset (Diverse Dermatology Images, DDI) and a disease-shifted tone-diverse dataset (Skin Condition Image Network, SCIN). Our results show that disease-distribution shift contributes more than skin tone in the evaluated settings. The cancer baseline decreases from 0.62 to 0.21 balanced accuracy when transferred to unfamiliar clinical conditions, while the within-disease skin-tone gap is smaller (0.10-0.18) and inconsistent. Label-free representation analysis shows that this failure reflects a representational limitation rather than only missing output labels: cancer-specialized features poorly cluster unfamiliar conditions (kNN purity lift +0.06 over chance), whereas dermatology-pretrained features retain stronger transferable structure (+0.23). Finally, we show that representation quality predicts recoverable performance under lightweight adaptation. Starting from dermatology foundation models, approximately ten labeled examples per clinical category recover most attainable performance. We release the evaluation protocol and code to support reproducible auditing of dermatology AI generalization.