FairLens:面向高风险决策的视觉语言模型公平性基准测试
FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making
浏览论文内容
中文总结 AI 辅助
该研究推出FAIRLENS基准框架,评估8个VLMs在招聘、法律、医疗领域的公平性与有效性,发现其核心问题是无根据推理,而非不平等待遇,该框架可迁移至带人口统计注释的人脸语料库。
中文摘要 AI 辅助
视觉语言模型(VLMs)越来越多地用于基于视觉输入做出决策。我们推出FAIRLENS,这是一个用于衡量VLMs在招聘、法律和医疗三个高风险领域中响应的公平性和有效性的基准与评估框架。FAIRLENS将涵盖性别、种族和年龄群体的真实人脸图像与封闭和开放式问题配对,每个模型提供超过10万组图像-问题对,并从四个互补视角评估响应:不利结果率的人口统计 parity(均等性)、可靠性、无支持角色和身份的人口统计关联,以及自由文本生成中的偏差。可靠性是核心有效性标准:当响应遵循问题中陈述的证据且图像无法支持答案时,该响应即为可靠(此时需弃权(不执行))。我们对8个VLMs进行评估,发现主要失败点是无根据的推理而非不平等待遇。模型通常从人脸推断资质、威胁、疾病或职业角色,而非弃权,最弱的模型在其输入无法回答的问题中99%都会如此。这些失败在法律和医疗领域最为严重,而这些领域最需要识别证据不足,仅差异指标会忽略它们:parity(均等性)差距绝对值很小,但当基线不利率较低时,相同差距意味着某一人口群体获得不利标签的频率是另一群体的数倍,且小差距同样可能反映模型对每个群体的处理都不安全。自由文本响应中的偏差与多项选择题准确率仅松散关联,因此正确的结构化答案并不意味着生成安全。FAIRLENS表明,公平的高风险VLM行为需要跨群体的相似处理,且拒绝从外观推断高风险属性,其问题套件可迁移到任何带人口统计注释的人脸语料库。
英文摘要
Vision-language models (VLMs) are increasingly used to make decisions from visual inputs. We introduce FAIRLENS, a benchmark and evaluation framework for measuring both the fairness and the validity of VLM responses in three high-stakes domains: hiring, legal, and healthcare. FAIRLENS pairs real face images spanning gender, race, and age groups with closed- and open-ended questions, giving more than 100K image-question pairs per model, and evaluates responses from four complementary views: demographic parity over adverse outcome rates, soundness, demographic association over unsupported roles and statuses, and bias in free-text generation. Soundness is the central validity criterion: a response is sound when it follows the evidence stated in the question and abstains when the image cannot support an answer. Evaluating eight VLMs, we find that the primary failure is unwarranted inference rather than unequal treatment. Models routinely infer qualifications, threat, illness, or professional role from a face instead of abstaining, and the weakest model does so on 99% of the questions its input cannot answer. These failures are most severe in legal and healthcare, where recognizing insufficient evidence matters most, and disparity metrics alone would miss them: parity gaps are small in absolute terms, yet when baseline adverse rates are low the same gap means one demographic group receives adverse labels several times as often as another, and a small gap can equally reflect a model that treats every group unsafely. Bias in free-text responses is only loosely coupled to multiple-choice accuracy, so correct structured answers do not imply safe generation. FAIRLENS shows that fair high-stakes VLM behavior requires similar treatment across groups and refusal to infer high-stakes attributes from appearance, and its question suite transfers to any face corpus with demographic annotations.
发表机构
- Vector Institute(矢量研究院)
机构由 AI 辅助整理,请以论文原文为准。