算法性别预测是不合法的,但性别推断可产生有效测量结果
Algorithmic Gender Prediction Is Illegitimate, But Gender Imputation Can Yield Valid Measurements
AI总结:
该研究区分了性别预测的合法性与有效性,指出用于公平的性别推断对部分性别歧视场景可产生有效测量但对跨性别者等不合法,建议仅在必要时部署并开发更包容的方法。
AI中文摘要:
机器学习伦理研究者和关键的人机交互(HCI)学者认为,算法化预测性别是错误的。与此同时,其他研究者依赖预测出的性别标签来研究性别差异、开发算法公平技术。我们如何调和这两种看似矛盾的直觉?我们区分了性别预测可能出错的两种方式:一是不合法,从而造成伤害;二是无效,从而产生无法使用的测量结果。我们将反对性别预测的论点转化为合法性与有效性的术语,表明用于公平目的的性别推断可能不合法,但仍能产生有效的差异测量结果。我们借助跨女性主义文献阐明这种矛盾,将针对女性与女性气质的性别歧视,与针对跨性别者和非二元性别者的性别歧视区分开来。虽然性别推断可对前者产生有效测量结果,但对后者而言它是不合法且有害的。我们主张从业者应仅在无法通过其他合理方式实现反歧视益处、且伤害被尽可能最小化时,才部署性别推断。我们通过三个案例研究考察这种张力:生成式图像模型中的性别偏差审计、电影中的性别差异测量、从个人姓名推断性别。通过将合法性与有效性区分开,并区分两种形式的性别歧视,我们表明关于性别预测的争论混淆了不同的关切,既模糊了性别推断可支持公平工作的场景,也模糊了其根本无法捕捉的针对跨性别者和非二元性别者的伤害。我们最后建议开发更具包容性的方法,以应对所有类型的性别歧视。
英文摘要:
Machine learning ethics researchers and critical HCI scholars have argued that algorithmically predicting gender is wrong. At the same time, other researchers rely on predicted gender labels to study gender disparities and develop algorithmic fairness techniques. How do we reconcile these two seemingly contradictory intuitions? We differentiate two ways gender prediction may be wrong: being illegitimate, thereby contributing to harm; and being invalid, thereby producing unusable measurements. Our analysis translates arguments against gender prediction into these terms of legitimacy and validity and shows how gender imputation applied for fairness purposes can be illegitimate yet still yield valid disparity measurements. We clarify this bind by drawing upon transfeminist literature to distinguish sexism that targets women and femininity from sexism that targets transgender and nonbinary people. While gender imputation can produce valid measurements for the former, it is illegitimate and harmful for the latter. We argue that practitioners should deploy gender imputation only when it would achieve anti-discrimination benefits that cannot be achieved through other reasonable means, while harms are minimized to the extent possible. We examine this tension in three case studies: auditing gender bias in generative image models, measuring gender disparities in film, and imputing gender from personal names. By disentangling legitimacy from validity, and differentiating these two forms of sexism, we show how debates over gender prediction have conflated distinct concerns, obscuring both the settings in which gender imputation can support fairness efforts and the harms towards transgender and nonbinary people that it fundamentally cannot capture. We conclude by recommending the development of more inclusive methods that address all kinds of sexism.