稠密检索模型中性别敏感性的机制分析
A Mechanistic Analysis of Gender Sensitivity in Dense Retrieval Models
浏览论文内容
中文总结 AI 辅助
本文通过机制分析定位稠密检索模型性别敏感性源于输入嵌入与后期层注意力头,测试不同引导干预效果,为针对性去偏提供机制基础并指出相关信号分离挑战。
中文摘要 AI 辅助
尽管稠密检索模型中的性别偏见已被充分证实,现有研究表明这类模型通常对男性性别指向的文档打分高于女性或中性变体,但产生这些差异的内部机制却鲜为人知。本文对双编码器模型进行机制分析以定位性别敏感性,发现该信号源自输入嵌入,并通过一小部分携带性别与术语匹配信号的后期层注意力头传播。基于这些发现,我们在已识别的两个节点处测试了引导干预,发现效果存在差异:嵌入层面的引导会非特异性地中和分数差异,而注意力层面的引导则会产生定向偏移。我们的发现为针对性去偏提供了机制基础,并凸显了在共享模型组件中分离性别与相关性信号的挑战。
英文摘要
While gender bias in dense retrieval models is well documented, with prior work showing that models often score male-gendered documents higher than female or neutral variants, the internal mechanisms producing these disparities are poorly understood. In this paper, we mechanistically analyze bi-encoder models to localize gender sensitivity, finding that the signal originates in input embeddings and propagates through a small set of late-layer attention heads that carry both gender and term-matching signals. Guided by these findings, we test steering interventions at both identified points and find distinct effects: embedding-level steering non-specifically neutralizes score differences, while attention-level steering produces directional shifts. Our findings provide a mechanistic basis for targeted debiasing and highlight the challenge of disentangling gender from relevance signals in shared model components.