arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16651cs.MM

视觉-语言模型的机制级评估:性别偏见的受控激活替换诊断

Mechanism-Level Evaluation for Vision-Language Models: Controlled Activation-Replacement Diagnosis of Gender Bias

Zhipeng Zhao, Wenxu Wang, Peishun Liu, Ruichun Tang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出机制级评估,通过因果中介分析诊断视觉-语言模型中的性别偏见,揭示层间敏感性模式,并指出其与干预效果弱相关,需结合行为基准。

中文摘要 AI 辅助

行为基准测试揭示了视觉-语言模型中存在哪些偏见,但无法揭示哪些内部组件对定向干预最敏感,从而阻碍了有原则的干预。我们认为机制级评估是必要的补充,并论证因果中介分析可作为性别偏见的诊断工具。我们将性别线索效应分解为归因于特定层激活的受控间接效应和通过所有其他路径的直接效应,从而生成逐层的机制特征。在跨越三个架构家族(LLaVA-1.5、LLaVA-NeXT、InstructBLIP,规模为7B/13B)和两种8B规模架构的六个模型中,出现了三个发现:在受控干预下,语言层激活表现出最大的输出敏感性,且直接成分通常带有相反的符号;架构选择重新分配了层间对激活替换的敏感性;反事实分数与表面分数存在分歧,揭示了隐含关联。系统性消融验证了内部一致性。一项干预实验发现,平均间接效应(AIE)与下游干预有效性仅呈弱相关(Pearson r = 0.33),且AIE第二大的层产生的偏见变化接近于零——这表明机制诊断捕捉了激活替换敏感性,但本身并不能确定最佳干预目标。这些结果表明,机制级评估能够捕捉行为基准无法捕捉的架构特定敏感性模式;将两者结合应成为标准NLP实践。代码:此https URL。

英文摘要

Behavioral benchmarking reveals \emph{what} biases exist in vision-language models but not \emph{which internal components} are most sensitive to targeted intervention, precluding principled intervention. We argue for mechanism-level evaluation as a necessary complement, demonstrating causal mediation analysis as a diagnostic instrument for gender bias. We decompose gender-cue effects into controlled indirect effects attributable to specific-layer activations and direct effects through all other pathways, producing layer-by-layer mechanistic signatures. Across six models spanning three architectural families (LLaVA-1.5, LLaVA-NeXT, InstructBLIP at 7B/13B) and two 8B-scale architectures, three findings emerge: language-layer activations exhibit the greatest output sensitivity under controlled intervention, with the direct component often carrying the opposite sign; architectural choices redistribute layer-wise sensitivity to activation replacement; and counterfactual scores diverge from surface-level scores, exposing implicit associations. Systematic ablation validates internal consistency. An intervention experiment finds that the average indirect effect (AIE) and downstream intervention effectiveness are only weakly correlated (Pearson $r = 0.33$), and the layer with the second-largest AIE produces near-zero bias change---indicating that mechanistic diagnosis captures activation-replacement sensitivity but does not, by itself, identify optimal intervention targets. These results show mechanism-level evaluation captures architecture-specific sensitivity patterns that behavioral benchmarks cannot; pairing both should become standard NLP practice. Code: https://github.com/zhaozhipeng1997/CARD-GenderBias.

发表机构

  • Ocean University of China(中国海洋大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑