arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ViD:面向大型视觉-语言模型的视觉主导性别偏见缓解方法

ViD: Vision-Dominant Gender Bias Mitigation for Large Vision-Language Models

Zhipeng Zhao, Zhaoqiang Wei, Peishun Liu, Youwei Zhao, Ruichun Tang

arXiv 2609.16647首次发表:更新:

发表机构

Ocean University of China(中国海洋大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ViD因果框架,通过分析注意力模式并采用后门调整与解码层令牌选择双重机制,在不增加训练开销下显著降低LVLMs性别偏见,提升视觉基础与输出公平性。

AI 中文摘要

大型视觉-语言模型(LVLMs)中的性别偏见损害了其公平性和可靠性,降低了输出可信度。现有的缓解方法依赖于训练阶段调整或事后校准,但在动态视觉偏见缓解方面存在局限,包括无法捕捉实时视觉-文本不一致性、依赖预定义的性别偏见分类体系,以及面对新出现的偏见模式时跨模态对齐性能下降。为解决这些问题,我们提出了ViD,一个受因果启发的框架,通过分析五种不同模式下的注意力机制,揭示了强语言先验带来的混杂效应。ViD证明,视觉到语言的交叉注意力能有效抑制偏见,同时保持通用推理能力和文本生成质量。ViD包含双重机制:后门调整用于对抗强语言先验,而解码层中的精细令牌选择则优化处理过程,从而增强模型鲁棒性和推理效率。我们的集成方法显著缓解了LVLMs中多维社会属性上的性别偏见,改善了视觉基础定位和输出公平性。跨基准验证表明,ViD在单属性评估(FACET)上将性别偏见降低了14.7%,并在图像描述任务(MS COCO)上取得了显著改进,其中LLaVA的性别偏见评分从0.6708提升至0.9978。关键在于,这些改进无需额外训练开销,使ViD成为LVLMs偏见缓解的可扩展且实用的解决方案。

英文摘要

Gender bias in large vision-language models (LVLMs) undermines their fairness and reliability, compromising output trustworthiness. Current mitigation methods rely on training-phase adjustments or post-hoc calibration, but face limitations in dynamic visual bias mitigation. These include inability to capture real-time visual-textual incongruence, dependence on predefined gender bias taxonomies, and degraded cross-modal alignment with emergent bias patterns. To address these challenges, we propose ViD, a causally-inspired framework that analyzes attention mechanisms across five distinct patterns, revealing confounding effects from strong language priors. ViD demonstrates that visual-to-language cross-attention effectively suppresses bias while preserving general reasoning capabilities and text generation quality. ViD incorporates dual mechanisms: backdoor adjustment counters strong language priors, while refined token selection in decoding layers optimizes processing. This enhances model robustness and inference efficiency. Our integrated approach significantly mitigates gender bias across multidimensional social attributes in LVLMs, improving visual grounding and output fairness. Cross-benchmark validation shows ViD reduces gender bias by 14.7\% on single-attribute evaluations (FACET) and achieves significant improvements on image captioning tasks (MS COCO), with gender bias score improving from 0.6708 to 0.9978 for LLaVA. Crucially, these improvements require no additional training overhead, making ViD a scalable and practical solution for bias mitigation in LVLMs.

CommentsEMNLP 2026 Main

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑