胸部X射线基础模型适配策略的子组性能分析
Subgroup performance analysis of adaptation strategies for chest X-ray foundation models
另 1 家 · 查看机构详情
- Imperial College London(帝国理工学院)
- Royal Brompton Hospital(皇家布朗普顿医院)
- Causality in Healthcare AI Hub(医疗保健AI因果关系中心)
- National Heart & Lung Institute(国家心肺研究所)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对胸部X射线基础模型,探究三种参数高效适配技术对病理分类性能与子组公平性的影响,发现整体性能提升未必减少子组差异,公平性影响需直接按任务评估
中文摘要 AI 辅助
基础模型正越来越多地被适配用于下游医学影像任务,但所选适配策略对子组公平性的影响仍鲜为人知。我们研究了三种参数高效的适配技术——包括在原始<[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]> token上的线性头、多层感知器(MLP)以及对多层patch特征的注意力池化模块——当应用于冻结的Rad-DINO胸部X射线编码器时,如何影响病理分类性能和子组差异。我们使用MIMIC-CXR数据集,在保留患病率、人口统计学均衡的测试集上评估了8种病理在种族、性别和成像视图子组上的表现,还进一步探究了每个适配器对受保护属性的编码强度。我们发现,注意力池化实现了最强的整体判别性能,且对属性(尤其是种族)的编码最强,但整体性能提升并未始终减少子组差异。值得注意的是,更强的属性编码并不对应更大的差异:网络早期层对种族的编码最弱,却产生了最大的子组性能差距。探索不同的注意力池化层组合进一步发现,被池化的层、属性编码强度与子组公平性之间不存在一致的关系。我们的结果表明,更丰富、更具表达力的表示可提高准确性,而公平性影响则取决于任务且不可预测,必须直接针对每个任务进行评估,而非仅从编码强度或整体性能推断。
英文摘要
Foundation models are increasingly adapted for downstream medical imaging tasks, yet the influence of the chosen adaptation strategy on subgroup fairness remains poorly understood. We investigate how three parameter-efficient adaptation techniques, including linear heads on the raw CLS token, an MLP, and an attention-pooling module over multi-layer patch features, affect both pathology classification performance and subgroup disparities when applied to the frozen Rad-DINO chest X-ray encoder. Using MIMIC-CXR, we evaluate eight pathologies across race, sex, and imaging-view subgroups on a prevalence-preserving, demographically balanced test set, and additionally probe how strongly each adapter encodes protected attributes. We find that attention pooling achieves the strongest overall discriminative performance and encodes attributes, particularly race, most strongly, but that improved overall performance does not consistently reduce subgroup disparities. Notably, stronger attribute encoding did not correspond to larger disparities: early network layers encoded race most weakly yet produced the largest subgroup performance gaps. Exploring different attention-pooling layer combinations further revealed no consistent relationship between the layers pooled, attribute encoding strength, and subgroup fairness. Our results indicate that richer, more expressive representations can improve accuracy while leaving fairness implications task-dependent and unpredictable, which must be assessed directly and per-task rather than inferred from encoding strength or overall performance alone.