发表机构
KU Leuven; Augment, imec research group at KU Leuven; University of Glasgow(荷语鲁汶大学; 荷语鲁汶大学 Augment 研究所; 格拉斯哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过FairAware工具评估非技术用户在公平性评估中的理解,发现用户能识别偏见但存在主客观理解差距,需内置检查以支持高风险决策。
AI 中文摘要
公平性指标的选择通常由数据科学家负责,但哪些偏见存在问题以及哪种指标最能捕捉这些偏见,取决于利益相关者的经验和领域知识。这要求非技术利益相关者的参与,但迄今为止为此目的构建的研究原型尚未测试这些利益相关者是否对其交互的指标形成准确的思维模型,或能否利用这些模型来识别偏见。我们提出了FairAware,一个与人力资源(HR)领域专家共同设计的公平性评估工具。我们通过一项混合方法研究评估了利益相关者的理解,该研究涉及70名参与者(35名HR员工,35名求职者),测量了客观和主观理解、认知负荷、偏见识别准确性以及开放式反馈。大多数参与者正确识别了最不利的群体,任务持续时间是唯一显著的预测因素。我们还发现主观理解与客观理解之间存在差距,两组在所有测量中的表现相似。这些结果表明,面向非专家的公平性评估工具可用于识别偏见,但在利益相关者做出更高风险的决策之前,需要内置理解检查。
英文摘要
Fairness metric selection is typically left to data scientists, but which biases are problematic and which metric captures them best depends on stakeholders' experience and domain knowledge. This calls for involving non-technical stakeholders, but the research prototypes built for this purpose so far have not tested whether these stakeholders form accurate mental models of the metrics they interact with or can act on them to identify biases. We present FairAware, a fairness assessment tool co-designed with Human Resources (HR) domain experts. We evaluate stakeholders' understanding through a mixed-methods study with 70 participants (35 HR employees, 35 job seekers), measuring objective and subjective understanding, cognitive load, bias identification accuracy, and open-ended feedback. Most participants correctly identified the most disadvantaged group, with task duration being the only significant predictor. We also found a gap between subjective and objective understanding, with both groups performing similarly across all measures. These results suggest that fairness assessment tools for non-experts are usable for identifying biases but need built-in checks on understanding before stakeholders make higher-stakes decisions.
Comments10 pages and 7 figures, excluding references and appendices; 26 pages and 10 figures in total. To be published in AIES 2026's proceedings