arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度偏差:对大型视觉语言模型中社会偏差的自适应深度探测

DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs

Anqi Li, Jie Zhang, Zhongqi Wang, Songkai Xue, Jiahao Wang, Shiguang Shan, Xilin Chen

arXiv 2607.11228首次发表:更新:

AI 中文总结

研究大型视觉语言模型中社会偏差问题,提出DeepBias自适应框架,通过“生成-演化-探测”循环及智能体深度探测偏差,构建DeepBiasBench基准,实验证明其有效性,为LVLM安全评估建立进化范式。

AI 中文摘要

虽然大型视觉语言模型(LVLMs)展现出卓越能力,但极易受到内在社会偏差影响。现有偏差评估协议主要依赖静态数据集,只能提供表面评估。我们引入DeepBias,一个通过精心设计的智能体对LVLMs中社会偏差进行深度探测的自适应框架。其通过动态的“生成-演化-探测”循环运行。首先,生成提议智能体合成测试数据并基于目标LVLMs的响应通过直接偏好优化(DPO)迭代更新,探索模型特定的失败模式。其次,自主技能驱动的挖掘智能体在多个探测回合中重写每个测试数据,从精心策划的深化和重写策略技能库中自适应选择。每回合该过程以模型先前响应为条件,能暴露更深层次的偏差。此外,我们用该框架构建了名为DeepBiasBench的基准。通过使用五个不同的先进LVLMs集合作为锚点,该基准捕获了跨架构共享的漏洞。综合实验证明了我们框架的有效性,表明DeepBias为深度偏差评估提供了具有挑战性的基准,为LVLM安全评估建立了进化范式。

英文摘要

While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predominantly rely on static datasets, which provide only a superficial assessment, as their fixed test cases cannot adaptively evolve to measure the true depth and limits of model vulnerabilities. We introduce DeepBias, an adaptive framework for the in-depth probing of social biases in LVLMs with carefully designed agents. Our approach operates through a dynamic ''generation-evolution-probing'' loop. First, a generative ProposerAgent synthesizes test data and is iteratively updated via Direct Preference Optimization (DPO) based on the target LVLM's responses, exploring model-specific failure modes. Second, an autonomous skill-driven DiggerAgent rewrites each test data across multiple probing turns, adaptively selecting from a curated skill library of deepening and rewriting strategies. At each turn, this process is conditioned on the model's previous response, enabling progressively deeper biases to be exposed. Furthermore, we build a benchmark named DeepBiasBench using our framework. By employing an ensemble of five diverse state-of-the-art LVLMs as anchors, the benchmark captures vulnerabilities shared across architectures. Comprehensive experiments demonstrate the effectiveness of our framework and show that DeepBias provides a challenging benchmark for in-depth bias evaluation, establishing an evolutionary paradigm for LVLM safety assessment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑