arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28014math.STstat.TH

OT-FairBoost:面向表格数据公平性正则化的最优传输引导梯度提升

OT-FairBoost: Optimal Transport-Guided Gradient Boosting for Fairness Regularization on Tabular Data

Veronika Shilova, Abdoulaye Sakho, Younes Boumoussou, Laurent Risser, Jean-Michel Loubes, Emmanuel Malherbe

中文总结 AI 辅助

本研究提出新型在处理框架OT-FairBoost,将Wasserstein-2距离惩罚融入梯度提升树目标函数,集成于LightGBM训练流程,在多任务场景下实现最优准确率-公平性权衡。

中文摘要 AI 辅助

尽管基于神经网络的机器学习模型近来备受关注,但梯度提升等树模型在表格数据任务中仍具竞争力,因此在各类AI应用中被广泛使用。与其他机器学习预测模型类似,它们可能因算法偏见在不同人口统计群体间产生歧视性预测,这类不良现象催生了新的监管框架和多种AI公平性策略。目前虽有多种预处理和后处理方法可缓解梯度提升模型的此类偏见,但仅少数在处理方法被提出。为填补这一空白,我们引入OT-FairBoost,这是一种将Wasserstein-2距离惩罚直接融入梯度提升树目标函数的新型在处理框架。这种基于OT的缓解策略已被证明能有效优化神经网络预测的群体公平性准则,如统计 parity(人口 parity)和均等机会。为将该方法适配梯度提升,我们将群体预测间Wasserstein-2距离的样本级梯度估计扩展至离散分布和海森对角线,随后将该方法集成到LightGBM训练流程中,并在二分类、回归及多群体敏感属性设置下进行评估。各设置下的实验结果表明,OT-FairBoost相较于其他方法实现了最优的准确率-公平性权衡。

英文摘要

Although neural-based machine learning models have received a lot of attention recently, tree-based models such as gradient boosting are competitive for tabular data and therefore remain widely used in various applications of AI. As when using other machine learning predictive models, they can however yield discriminative predictions across demographic groups, due to so-called algorithmic biases. These undesirable phenomena have motivated the emergence of new regulatory frameworks and various AI fairness strategies. While several pre-and post-processing methodologies exist to mitigate such bias on gradient boosting models, only a few in-processing methods have been proposed. To bridge this gap, we introduce OT-FairBoost, a novel in-processing framework that incorporates a Wasserstein-2 distance penalty directly into the objective function of gradient-boosted trees. This OT-based mitigation strategy has been shown to efficiently optimize group fairness criteria such as Demographic Parity and Equalized Odds on neural-based predictions. To adapt this approach for gradient boosting, we extend the sample-wise gradient estimation of the Wasserstein-2 distance between group predictions to discrete distributions and hessian diagonals. We then integrate our approach into the LightGBM training procedure and evaluate it across binary classification, regression, and multi-group sensitive attribute settings. Experimental results in each of these settings demonstrate that OT-FairBoost achieves best accuracy-fairness trade-offs against alternatives.

↑