使用PCA识别数据集中的表征偏差:一种最大差异划分框架
Identifying Representational Biases in Datasets Using PCA: A Max-Disparity Partition Framework
浏览论文内容
中文总结 AI 辅助
本研究提出一种基于Fiduccia-Mattheyses框架的贪心局部搜索算法,在无组标签情况下通过PCA识别数据集中最大表征差异的划分,并在学生数据集上验证了检测-解释-缓解流程。
中文摘要 AI 辅助
主成分分析(PCA)最小化总体重构误差,这可能会无意中使多数子群体以远高于少数子群体的保真度被表征。PCA的公平性感知扩展纠正了这种差异,但需要组标签作为输入。我们处理一个逻辑上更前置的问题:仅给定一个数据矩阵,在共享的PCA投影下,数据的哪个二分类划分遭受最大的表征差异?我们将此形式化为最大差异划分问题,并提出一种基于Fiduccia-Mattheyses二分划分框架的贪心局部搜索算法,该算法无需任何预定义的组标签即可发现差异最大化的划分。两种基准算法,即固定投影排序基线和模拟退火变体,证实了贪心解在经验上接近最优。在识别出划分后,我们通过PCA载荷得分和关联规则挖掘将差异归因于特定特征,使从业者能够评估弱势群体是否对应于人可理解的少数群体。在“预测学生辍学与学业成功”数据集上,表征差异主要由社会经济劣势的制度性和程序性代理指标驱动,而性别在弱势群体中作为次要但一致的贡献因素出现。发现的划分随后直接传递给公平PCA,完成了一个检测-解释-缓解流程。
英文摘要
Principal Component Analysis (PCA) minimises aggregate reconstruction error, which can inadvertently represent majority subgroups with substantially higher fidelity than minority subgroups. Fairness-aware extensions of PCA correct this disparity but require group labels as input. We address the logically prior question: given only a data matrix, which binary partition of the data suffers the greatest representational disparity under a shared PCA projection? We formalise this as the max-disparity partition problem and propose a greedy local-search algorithm, grounded in the Fiduccia-Mattheyses bipartitioning framework, that discovers the disparity-maximising partition without any predefined group labels. Two benchmark algorithms, a fixed-projection sorting baseline and a simulated-annealing variant, confirm that the greedy solution is empirically near-optimal. Having identified the partition, we attribute the disparity to specific features via PCA loading scores and association rule mining, enabling a practitioner to assess whether the disadvantaged group corresponds to a human-meaningful minority. On the Predict Students' Dropout and Academic Success dataset, representational disparity is driven predominantly by institutional and programmatic proxies for socioeconomic disadvantage, with gender emerging as a secondary but consistent contributor within the disadvantaged group. The discovered partition is then passed directly to Fair PCA, completing a detect-explain-mitigate pipeline.
发表机构
- Indian Institute of Science(印度科学学院)
机构由 AI 辅助整理,请以论文原文为准。