发表机构
State Key Lab of Software Development Environment, Beihang University(北京航空航天大学软件开发环境国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大视觉语言模型的社会偏差问题,提出反事实集成解码框架,通过构建多群体反事实视角并集成解码,在三类基准上实现最高47.97%的偏差降低,同时保留模型核心能力。
AI 中文摘要
大视觉语言模型(LVLMs)在各类任务中已取得显著性能,但它们常从训练数据中继承社会偏差,导致处理不同社会群体肖像时出现有偏差的行为。现有去偏差方法通常在解码时比较原始生成与有偏差生成的标记概率,但这些方法根本上受限于依赖单一刻板观点,未能考虑社会视角的多样性。受社会科学中“多样性促进公平”原则启发,我们提出Counterfactual Ensemble Decoding(CED,反事实集成解码)这一新颖框架,该框架在视觉表示空间内构建多群体反事实视角,并在解码过程中整合这些视角以促进公平的模型行为。CED首先在视觉空间中执行反事实引导,方法是识别与每个社会群体相关的语义方向,并沿这些方向生成反事实表示,从而提供破坏刻板叙事的多样化视角。在解码过程中,CED定位这些视角间分歧最大的解码器层,并使用感知不确定性的权重集成它们的标记分布,优先考虑不同群体的高置信度标记,以产生更均衡的概率分布,指导更公平的生成。在三个社会偏差评估基准上进行的大量实验表明,所提方法相较于领先基线取得了显著改进,在涉及职业、描述符和人格特质的场景中,偏差降低幅度最高达47.97%。此外,CED还以最小的性能下降保留了原始模型的核心能力。
英文摘要
Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases from their training data, resulting in biased behavior when processing portraits from different social groups. Existing debiasing approaches typically compare token probabilities between the original and biased generations during decoding, but they are fundamentally limited by their reliance on a single, stereotyped viewpoint and fail to account for the diversity of social perspectives. Inspired by the social science principle that diversity fosters fairness, we propose Counterfactual Ensemble Decoding (CED), a novel framework that constructs multi-group counterfactual perspectives within the visual representation space and integrates them during decoding to promote equitable model behavior. CED first performs counterfactual steering in the visual space by identifying semantic directions associated with each social group and generating counterfactual representations along these directions, thereby offering diverse perspectives that disrupt stereotypical narratives. During decoding, CED locates the decoder layer exhibiting the greatest divergence among these perspectives and ensembles their token distributions using uncertainty-aware weights, prioritizing high-confidence tokens from different groups to yield a more balanced probability distribution that guides fairer generation. Extensive experiments on three social bias evaluation benchmarks demonstrate that \tool achieves substantial improvements over leading baselines, reducing bias by up to 47.97% across scenarios involving occupations, descriptors, and persona traits. Moreover, CED also preserves the core capabilities of the original model with minimal degradation.