发表机构
Friedrich-Alexander-Universität Erlangen-Nürnberg; RWTH Aachen University; University Hospital RWTH Aachen(埃尔朗根-纽伦堡弗里德里希-亚历山大大学; 亚琛工业大学; 亚琛工业大学医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FRAME是用于医学影像公平性审计的两步框架,可分离采样变异与表征成因,能解释部分性能差异,辅助区分需机制解释的差异与采样变异可解释的差异。
AI 中文摘要
亚组性能差异是医学影像公平性偏差的标准证据,常规应对方式是移除模型编码的人口统计信息。本文提出Fair-model Reference And Mechanism Evaluation(FRAME,公平模型参考与机制评估),这是用于验证该主张的两步框架。第一步推导公平模型参考,即观测亚组规模下完全公平时的差异分布;第二步在表征空间中用两个算子检验剩余部分,其中一个算子按构造无法改变组内排序。在702206张图像和36个编码器上,该参考可解释中位数41%的报告种族差异和22%的年龄差异。注入人口统计可解码性不会改变剩余差异,而将群体与疾病方向纠缠会使种族差异从0.077升至0.118。测试的所有干预措施对剩余差异的影响均不超过随机种子变化的影响,这些干预措施在操作点减小差异,但使组内排序差异保持在中位数0.000。将FRAME应用于6种医学影像模态的9项已发表研究中的89个差异时,该参考可解释中位数25%的率差异和70%的受试者工作特征曲线下面积差异;而图像-文本预训练会使最差组性能提升约0.05。在选择干预措施前应用FRAME,可区分需要机制解释的差异与当前队列规模下采样变异可解释的差异。
英文摘要
Fairness audits of medical imaging models commonly report the largest performance difference between demographic subgroups and treat any positive value as evidence of bias. Yet a perfectly fair model also produces a positive difference, which grows as subgroups shrink. Some mitigation methods remove demographic information from models because they assume that it produces the difference. Here we introduce Fair-model Reference And Mechanism Evaluation (FRAME). Its first step compares each reported difference with the difference expected from a perfectly fair model at the same subgroup counts. Its second step tests whether candidate causes of any excess over this fair-model reference change the difference when they are injected into model features. We evaluated 36 encoders on nine datasets in three modalities and audited 89 published differences across six modalities. For 10 frozen encoders and 13 chest radiograph findings, the reference is a median 41% of the race difference and 22% of the age difference. It is exceeded by 40 of 53 published differences in sensitivity or false positive rate and by one of 36 in the area under the receiver operating characteristic curve (AUROC). Adding race information to the features of two encoders does not change the race difference significantly. At matched disease performance, mitigation reduces the race difference by less than its spread across three pretraining seeds (medians 0.005 and 0.012) and the age difference by more (0.008 and 0.006). Image-text pretraining raises the worst-group AUROC for race by a median 0.049 over self-supervised pretraining. Reporting the fair-model reference beside every new and published subgroup difference could direct mitigation to the differences above it.