arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

公平剪枝:通过差分激活定位GLU-MLP层中的人口统计偏差

Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

Pere Martra, Eugenio Martínez Cámara, Alfonso Ureña López

arXiv 2607.28319首次发表:更新:

AI 中文总结

本研究提出Fairness Pruning方法,通过最小对比提示对与激活捕获定位LLM的GLU-MLP层中人口统计偏差神经元,经实验验证该方法可针对性调控偏差且对模型能力影响极小。

AI 中文摘要

本研究提出了Fairness Pruning,一种轻量级结构干预方法,旨在管理和未来缓解大型语言模型(LLM)中的人口统计偏差。作为该方法的基础实证验证,本研究聚焦于因果偏差定位。通过使用最小对比提示对和推理时激活捕获,该方法识别GLU架构中处理人口统计属性时反应存在差异的神经元,在down_proj输入处评估信号。实证评估在参数规模最高达30亿的模型(Llama-3.2系列和Salamandra-2B)上开展,结合标准化基准评估与定性文本生成实验。结果表明,将识别出的神经元归零会改变模型对相关人口统计变量的响应,但这种干预并非实现平稳缓解,而是引发双向偏差不稳定:由于BiasScore为无符号值,候选集混合了推动和反对刻板印象的神经元,对总体偏差的净效应取决于哪种符号占主导。该干预极具针对性:在Llama-3.2-1B中最多归零40个神经元(占总MLP宽度的比例不足0.031%),可实现推理和通用知识能力的平均保留率达99.49%。这些发现实证证实人口统计偏差处理与模型能力在可分离的回路上运行,为从盲目归零转向定向行为调节奠定了方法论基础。

英文摘要

This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and future mitigation of demographic bias in large language models (LLMs). As a foundational empirical validation of this method, this work focuses on causal bias localization. Using minimally contrastive prompt pairs and inference-time activation capture, the method identifies neurons that react differentially when processing demographic attributes in GLU architectures, evaluating the signal at the down_proj input. Empirical evaluation was conducted on models of up to 3 billion parameters (Llama-3.2 family and Salamandra-2B), combining standardized benchmark evaluation with qualitative text generation experiments. Results demonstrate that zeroing the identified neurons alters how the model responds to associated demographic variables. However, rather than producing flat mitigation, the intervention causes bidirectional bias destabilization: because BiasScore is unsigned, candidate sets mix neurons that push toward and against the stereotype, and the net effect on aggregate bias depends on which sign dominates. The intervention is extremely surgical: zeroing at most 40 neurons in Llama-3.2-1B (less than 0.031% of total MLP width) achieves a mean retention of 99.49% in reasoning and general knowledge capabilities. These findings empirically confirm that demographic bias processing and model capabilities operate on dissociable circuits, establishing the methodological foundations for transitioning from blind zeroing toward directional behavior modulation.

Comments15 pages, 3 figures, 9 tables. Code and datasets publicly available

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑