用于公平高效恶性肿瘤分类的可控多样皮肤病图像生成
Controllable Generation of Diverse Dermatological Imagery for Fair and Efficient Malignancy Classification
- University of California, Santa Cruz(加州大学圣克鲁兹分校)
- University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对皮肤病诊断中缺乏多样标注图像的问题,提出cgDDI框架,能合成健康皮肤样本、映射罕见病变,支持自动分割掩码,通过多数据集验证,提高了恶性肿瘤分类准确率,公开相关资源推动公平性研究。
AI中文摘要:
准确的皮肤病诊断需要在不同人群中表现公平,但缺乏专业注释图像阻碍了进展。我们引入了cgDDI框架,它能合成逼真的健康皮肤样本,非参数地将单样本罕见病变映射到新肤色和位置,用最少10个训练样本进行高效参数生成。该框架支持人工和自动分割掩码,可扩展到无预制病变掩码的数据集。通过两个数据集验证,在DDI基准上,仅合成训练时恶性肿瘤分类准确率达86.4%,真实数据微调后达90.9%,跨数据集实验在F17k数据上准确率提高13.9%。我们还公开了相关图像、代码和生成模型。
英文摘要:
Accurate dermatological diagnosis naturally necessitates equitable performance across diverse populations, yet a systematic lack of expertly annotated images, especially for underrepresented skin tones and rare diseases, impedes progress toward measurably fair methods. We introduce cgDDI (Controllable Generation of Diverse Dermatological Imagery), a hybrid framework that (1) synthesizes realistic healthy skin samples without disturbing other input properties, (2) maps single-sample rare lesions onto novel skin-tones and locations non-parametrically, and (3) allows for efficient parametric generation with as few as 10 training samples. The framework supports both human and automated segmentation masking, enabling scalability to datasets without pre-made lesion masks. We grow a 656-image dataset by more than 400x and validate across two datasets: biopsy-confirmed Diverse Dermatology Images (DDI) and expert-verified Fitzpatrick17k (F17k). On the DDI benchmark, we achieve malignancy classification accuracy of 86.4% under synthetic-only training and 90.9% state-of-the-art performance with real data fine-tuning, alongside leading fairness metrics. Cross-dataset experiments show +13.9% accuracy improvements on unseen F17k data despite minimal disease overlap. We openly release 266k+ synthetic images, code, and generative models to further support fairness research at https://github.com/hectorcarrion/ControllableGenDDI.