arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29590cs.CV

强安全护栏下大型视觉语言模型的护栏无关社会偏见评估

Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language Models

Yusuke Hirota, Michael Ross Boone, Arun George Zachariah, Jibin Rajan Varghese, Yu-Chiang Frank Wang, Boyi Li, Ryo Hachiuma

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对强护栏LVLMs的偏见评估难题,提出解耦人物属性推断的评估方法,发现20款LVLMs均存在不当使用用户人口统计信息的偏见,且专有模型偏见低于开源模型。

中文摘要 AI 辅助

我们提出了一种在强安全护栏时代针对大型视觉语言模型(LVLMs)的社会偏见评估方法。现有基准依赖于要求模型推断图像中人物属性的提示(例如“这个人是CEO还是秘书?”)。然而,我们发现具有强护栏的LVLMs,如GPT和Claude,经常会拒绝这些提示,导致评估不可靠。为解决这一问题,我们改变了先前的评估范式,将任务与所描绘的人物解耦:不再推断人物属性,而是使用不询问人物的提示(例如“写一个关于虚构人物的故事”),并将图像作为临时用户信息附上,以隐式提供人口统计线索,然后比较不同用户人口统计群体的输出。该方法在故事生成、术语解释和考试式问答三项任务中均得到实例化,即使在有护栏的LVLMs中也能避免拒绝,从而实现可靠的偏见测量。将其应用于20款近期的LVLMs(包括开源和专有模型),我们发现所有模型在与人物无关的任务中都不当使用了用户人口统计信息;例如,故事中的角色常被描绘为男性用户对应的机械师、女性用户对应的护士。尽管仍存在偏见,GPT-5这类专有模型的偏见程度低于开源模型。我们分析了这种差距背后的潜在因素,讨论了持续的模型监控与改进可能是减少偏见的一个原因。

英文摘要

We propose a societal bias evaluation method for large vision-language models (LVLMs) in the era of strong safety guardrails. Existing benchmarks rely on prompts that ask models to infer attributes of people in images (e.g., "Is this person a CEO or a secretary?"). However, we find that LVLMs with strong guardrails, such as GPT and Claude, often refuse these prompts, making evaluations unreliable. To address this, we change the prior evaluation paradigm by decoupling the task from the depicted person: instead of inferring person's attributes, we use prompts that do not ask about the person (e.g., "Write a fictional story about an imaginary person.") and attach the image as provisional user information to implicitly provide demographic cues, then compare outputs across user demographics. Instantiated across three tasks --- story generation, term explanation, and exam-style QA --- our method avoids refusals even in guardrailed LVLMs, enabling reliable bias measurement. Applying it to 20 recent LVLMs, both open-source and proprietary, we find that all models undesirably use user demographic information in person-irrelevant tasks; for instance, characters in stories are often portrayed as mechanic for male users and nurse for female users. Although still biased, proprietary models like GPT-5 show lower bias than open-source ones. We analyze potential factors behind this gap, discussing continuous model monitoring and improvement as a possible contributor for reducing bias.

发表机构

  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑