arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于偏差探针的多群体公平性高效主动审计

Efficient Active Auditing of Multi-Group Fairness with Bias Probes

Ayoub Ajarra, Debabrota Basu

arXiv 2609.40034首次发表:更新:

发表机构

Équipe Scool, Univ. Lille, Inria, CNRS, Centrale Lille, UMR 9189- CRIStAL(里尔大学、法国国家信息与自动化研究所、法国国家科学研究中心、里尔中央理工学院、CRIStAL实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出偏差探针框架及主动审计器ALeBi,在不重建模型的情况下高效估计多群体公平性指标,并揭示模型机密性与可靠审计间的权衡。

AI 中文摘要

在过去十年中,机器学习(ML)一直在双重目标下进行训练:通过经验风险最小化(ERM)最小化预测误差,同时控制不公平偏差。然而,在实践中,公平感知训练往往比标准ERM带来的改进有限,这使得可靠的事后审计变得至关重要。现有的黑盒模型审计方法要么依赖于模型重建——使系统暴露于提取攻击之下——要么直接估计公平性指标,对数据分布的哪些区域驱动偏差提供的洞察有限。更根本的是,针对特定属性的审计——旨在仅提取目标公平性信息而不重建模型——仍然知之甚少。在这项工作中,我们引入了偏差探针框架,该框架能够实现有针对性和自适应的查询以揭示偏差结构,同时保护模型机密性。基于这一框架,我们提出了ALeBi,一种主动审计器,它学习此类探针以高效估计多群体公平性指标。我们建立了由特定属性复杂度度量控制的新样本复杂度保证,解决了一个先前提出的开放问题,并将我们的分析扩展到模型所有者可能策略性地掩盖偏差的对抗性设置。我们的结果揭示了模型机密性与可靠审计之间的基本权衡,并表明特定属性探针能够实现准确估计以及高偏差和低偏差区域的可解释识别。大量实验支持了我们的理论发现,并证明了我们方法的实际有效性。

英文摘要

Over the past decade, Machine Learning (ML) has been trained under dual objectives: minimizing prediction error via Empirical Risk Minimization (ERM) while controlling unfairness bias. In practice, however, fairness-aware training often yields limited improvements over standard ERM, making reliable post hoc auditing essential. Existing auditing approaches for black-box models either rely on model reconstruction --exposing systems to extraction attacks-- or directly estimate fairness metrics, offering limited insight into which regions of the data distribution drive bias. More fundamentally, property-specific auditing --aimed at extracting only targeted fairness information without reconstructing the model-- remains poorly understood. In this work, we introduce the bias probe framework, which enables targeted and adaptive querying to reveal bias structure while preserving model confidentiality. Building on this framework, we propose ALeBi, an active auditor that learns such probes to efficiently estimate multi-group fairness metrics. We establish novel sample complexity guarantees governed by a property-specific complexity measure, resolving a previously posed open question, and extend our analysis to adversarial settings where the model owner may strategically obscure bias. Our results uncover a fundamental trade-off between model confidentiality and reliable auditing, and show that property-specific probing enables both accurate estimation and interpretable identification of high and low-bias regions. Extensive experiments support our theoretical findings and demonstrate the practical effectiveness of our approach.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑