发表机构
Hertie Institute for AI in Brain Health; University of Tübingen; Tübingen AI Center(赫蒂脑健康人工智能研究所; 蒂宾根大学; 蒂宾根人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出解释一致性得分(ECS)作为公平性感知指标,以糖尿病视网膜病变筛查为案例,发现不同族裔群体预测性能有差异但解释一致性较高且与性能差异无显著关联,说明预测公平性与解释一致性是模型行为的互补维度,推动公平性评估拓展至预测性能之外。
AI 中文摘要
医学成像领域的公平性通常通过子群体性能指标进行评估,但目前尚不清楚模型是否在不同人口统计群体中依赖一致的视觉证据。本研究引入了基于Jensen-Shannon散度的公平性感知指标——解释一致性得分(ECS),用于量化不同子群体间归因图的相似性。以糖尿病视网膜病变筛查为案例,在全局范围及疾病严重程度分层下对ECS进行评估。实验结果显示,尽管不同族裔群体的预测性能存在差异,但解释一致性仍保持相对较高水平,且与性能差异无显著关联。这些发现表明,预测公平性与解释一致性捕捉了模型行为的互补维度,为公平性评估需超越预测性能提供了依据。
英文摘要
Fairness in medical imaging is commonly evaluated through subgroup performance metrics, yet it remains unclear whether models rely on consistent visual evidence across demographic groups. This work introduces the Explanation Consistency Score (ECS), a fairness-aware metric based on Jensen-Shannon divergence that quantifies the similarity of attribution maps across subgroups. Using diabetic retinopathy screening as a case study, ECS is evaluated globally and within disease severity. Experiments reveal that while predictive performance differs across ethnic groups, explanation consistency remains relatively high and shows no significant association with performance disparities. These findings suggest that predictive fairness and explanation consistency capture complementary dimensions of model behavior, motivating fairness evaluations that extend beyond predictive performance.
CommentsAccepted for publication at the joint FAIMI, BRIDGE, and EPIMI Workshop at MICCAI 2026