发表机构
University of Illinois at Chicago; Pennsylvania State University(伊利诺伊大学芝加哥分校; 宾夕法尼亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出REMI框架,将反事实公平性视为关系不变量发现问题,可自动定位、解释和缓解个体歧视,在83%以上案例中定位真实公平性缺陷,使黑箱模型歧视性决策减少最多70%。
AI 中文摘要
数据驱动的软件系统越来越多地部署在刑事司法、金融借贷等高风险社会经济领域,但这些系统常表现出个体歧视,即程序仅因个体受保护属性(如种族、性别、年龄)不同,对相似个体产生不公平的结果差异。现有研究多聚焦于检测和量化这些缺陷,却严重缺乏用于解释和定位个体公平性缺陷的原则性机制;当前解释技术主要针对单输入决策设计,未考虑歧视的关系本质——其固有涉及原始样本与反事实样本对的比较。本文提出REMI框架,用于自动定位、解释和缓解个体歧视。该框架受形式方法中循环不变量合成的启发,将反事实公平性视为关系不变量发现问题;引入双向关系解释框架,在成对样本(x, x')上学习,以识别输入空间中公平性被违反的区域。与不变量推理中传统单向蕴含对不同,该方法强制执行双向约束:要求原始样本与反事实样本的结果一致。REMI利用三种数据对齐技术,推断可解释的基于规则的模型,即“公平性不变量”;这些规则作为护栏,可选择性阻止或重新标记不公平预测,无需重新训练模型。对符号程序和神经网络程序的评估显示,REMI在超过83%的案例中定位了真实公平性缺陷,显著优于现有基准方法,使黑箱模型的歧视性决策减少多达70%。
英文摘要
Data-driven software systems are increasingly deployed in high-stakes socio-economic domains, from criminal justice to financial lending. However, these systems often exhibit individual discrimination---unjustified disparities in which a program yields different outcomes for similar individuals who differ only in their protected attributes (e.g., race, gender, age). While existing research has focused on detecting and quantifying these bugs, there remains a critical lack of principled mechanisms to explain and localize individual fairness bugs. Current explanation techniques are largely designed for single-input decisions rather than the relational nature of discrimination, which inherently involves a comparison between an original and a counterfactual pair. We present REMI, a framework for the automated localization, explanation, and mitigation of individual discrimination. Inspired by loop-invariant synthesis in formal methods, we treat counterfactual fairness as a relational invariant discovery problem. We introduce a bidirectional relational explanation framework that learns over paired examples $(x, x')$ to identify regions of the input space where fairness is violated. Unlike traditional one-way implication pairs used in invariant inference, our approach enforces bidirectional constraints: requiring identical outcomes for both original and counterfactual samples. REMI utilizes three data-alignment techniques to infer interpretable rule-based models that act as "fairness invariants." These rules serve as guardrails to selectively block or relabel unfair predictions without requiring model retraining. Our evaluation on symbolic and neural network programs demonstrates that REMI localizes ground-truth fairness bugs in over 83% of cases, significantly outperforming state-of-the-art baselines and reducing discriminatory decisions in black-box models by up to 70%.
CommentsIn 35th edition of ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA 2026)