AI 中文总结
研究针对算法公平性方法多孤立单轴评估的局限,提出FairSelect工具包,支持多模型架构等评估。通过合成与真实临床数据集验证,发现公平干预相互作用方式复杂,该工具包能系统识别策略,兼顾临床机器学习中群体公平与模型性能。
AI 中文摘要
算法公平性方法越来越多地用于识别和减轻机器学习模型中的偏差,但大多数方法是孤立地沿着单一人口统计轴进行评估的。这限制了选择公平策略的实际指导,因为在交叉子群体和建模生命周期的多个阶段可能会出现差异。本文提出了FairSelect,这是一个用于系统评估公平性缓解策略的工具包,这些策略可在预处理、处理中和后处理阶段单独或组合应用。FairSelect支持多种模型架构、交叉子群体评估,以及跨基线、单一方法和多层次配置的公平性效用权衡比较。该框架使用旨在代表特定偏差机制的合成临床数据集以及心房颤动患者两年中风风险预测的真实世界复现进行了验证。合成实验表明,有针对性的公平性方法通常会减少预期的子群体差异,而组合策略在适度的效用权衡下产生了更大的平均公平性改善。在临床预测任务中,缓解效果差异很大,一些组合同时提高了公平性和预测性能,而另一些则无效或适得其反。这些发现表明,公平性干预以非加性和上下文相关的方式相互作用。FairSelect为系统识别公平性策略提供了一个实用框架,这些策略可在临床机器学习中提高子群体公平性,同时保持模型性能。
英文摘要
Algorithmic fairness methods are increasingly used to identify and mitigate bias in machine learning models, yet most approaches are evaluated in isolation and along single demographic axes. This limits practical guidance for selecting fairness strategies, where disparities may arise across intersectional subgroups and across multiple stages of the modeling lifecycle. This work presents FairSelect, a toolkit for systematically evaluating fairness mitigation strategies applied individually and in combination across preprocessing, inprocessing, and postprocessing stages. FairSelect supports multiple model architectures, intersectional subgroup evaluation, and comparison of fairness utility tradeoffs across baseline, single method, and multi level configurations. The framework was validated using synthetic clinical datasets designed to represent specific bias mechanisms and a real-world replication of two-year stroke risk prediction among patients with atrial fibrillation. Synthetic experiments showed that targeted fairness methods generally reduced intended subgroup disparities, while combined strategies produced larger average fairness improvements with modest utility tradeoffs. In the clinical prediction task, mitigation effects were highly variable, with some combinations improving both fairness and predictive performance while others were ineffective or counterproductive. These findings demonstrate that fairness interventions interact in nonadditive and context dependent ways. FairSelect provides a practical framework for systematically identifying fairness strategies that improve subgroup equity while preserving model performance in clinical machine learning.
Comments15 pages, 5 tables, Submission to Health Informatics Knowledge Management Conference 2026