arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于减轻疾病分类器偏差和检测偏差的人口统计学条件合成医学图像

Demographically-Conditioned Synthetic Medical Images for Bias Mitigation and Bias Detection in Disease Classifiers

Mahmoud Ibrahim, Bart Elen, Chang Sun, Gokhan Ertaylan, Michel Dumontier

arXiv 2607.14984首次发表:更新:

发表机构

Institute of Data Science, Faculty of Science and Engineering, Maastricht University; Department of Advanced Computing Sciences, Faculty of Science and Engineering, Maastricht University; VITO(马斯特里赫特大学科学与工程学院数据科学研究所; 马斯特里赫特大学科学与工程学院高级计算科学系; 比利时弗拉芒技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究医学图像分类器亚组公平性审计的样本量问题,提出用人口统计学条件合成生成器减轻训练偏差、检测评估偏差。以COVID-19胸部CT分类为例,发现其在偏差减轻和检测上的作用,如预训练效果好、能再现亚组排名等。

AI 中文摘要

医学图像分类器的每个亚组公平性审计面临样本量问题:在保留的测试集中,少数亚组的样本很少,以至于每个亚组性能的置信区间比审计旨在检测的偏差更宽。我们认为,人口统计学条件合成生成器可以同时做到:在训练方面减轻偏差,在评估方面检测偏差。通过使用端到端微调的Stable Diffusion 2.1生成器对COVID-19胸部CT分类进行研究,我们有两个发现。对于偏差减轻(训练),人口统计学平衡的合成队列作为预训练先验最有用,而不是作为联合增强:在相同的固定数据下,顺序预训练然后微调大大优于联合增强,并且得到的分类器在约100倍的真实数据效率下超过了全真实基线。对于偏差检测(评估),在五个合成少数队列和五个分类器种子中,合成估计器再现了强大的真实预言机的亚组排名(MCC和召回率上的Spearman ρ = 1.00),并在小的真实测试集样本耗尽的情况下给出更可靠的每个单元格估计。因此,合成队列在公平性审计关心的单元格中最有用,既是亚组偏差的修复方法,也是衡量方法。

英文摘要

Per-subgroup fairness audits of medical image classifiers face a sample-size problem: minority subgroups in held-out test sets have so few samples that the resulting confidence intervals on per-subgroup performance are wider than the bias the audit is meant to detect. We argue that a demographically-conditioned synthetic generator can do both: mitigate bias on the training side and detect bias on the evaluation side. Working on COVID-19 chest CT classification with an end-to-end fine-tuned Stable Diffusion 2.1 generator, we make two findings. For bias mitigation (training), a demographically-balanced synthetic cohort is most useful as a pretraining prior, not as joint augmentation: with the same fixed data, sequential pretraining followed by fine-tuning substantially outperforms joint augmentation, and the resulting classifier surpasses the full-real baseline at $\sim$$100\times$ real-data efficiency. For bias detection (evaluation), across five synthetic minority cohorts and five classifier seeds, the synthetic estimator reproduces the subgroup ranking of a well-powered real oracle (Spearman $ρ= 1.00$ on MCC and Recall) and gives the more reliable per-cell estimate where the small real test set runs out of samples. The synthetic cohort is therefore most useful in exactly the cells that fairness audits care about, as both a fix for and a measure of subgroup bias.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑