发表机构
Kyoto University; Zhejiang University; Shanghai Jiao Tong University; Inner Mongolia University(京都大学; 浙江大学; 上海交通大学; 内蒙古大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Acmite框架,通过概念引导和互信息最小化,实现对大语言模型性别偏见的选择性去偏,在保持通用能力的同时有效降低偏见。
AI 中文摘要
大型语言模型(LLMs)会从训练数据中再现社会刻板印象,这促使了关于模型去偏的广泛研究。然而,现有方法往往依赖于显式的偏见示例或预定义的群体术语替换,这使得它们对措辞敏感,并且难以捕捉跨不同语境共享的刻板印象概念。更重要的是,它们通常抑制有偏见的输出,而没有显式建模模型输出与潜在刻板印象概念之间的统计依赖性。我们提出了Acmite,一个轻量级的概念引导框架,用于有针对性和选择性的去偏。Acmite将刻板印象表示为结构化的语义概念,并使用最大边际相关性(MMR)来选择多样化的概念进行去偏。受互信息最小化的启发,它用token级别的KL散度来近似这种依赖性,同时保留任务语义。一个轻量级的LoRA适配器在冻结基础模型的情况下进行训练,并且仅在输入与刻板印象相关概念足够相似时在推理时激活;否则,直接使用原始模型。我们在BBQ、CrowS-Pairs和StereoSet上评估了Acmite,并在ARC-Challenge、GSM8K和PIQA上评估了通用能力的保持。跨三个LLM的实验表明,Acmite在互补的评估格式中有效减轻了性别偏见,同时在无关偏见的任务上保持了有竞争力的性能。匿名代码和数据可在该https URL获取。
英文摘要
Large language models (LLMs) can reproduce social stereotypes from their training data, motivating extensive research on model debiasing. However, existing methods often rely on explicit biased examples or predefined group-term substitutions, making them sensitive to wording and less effective at capturing stereotype concepts shared across diverse contexts. More importantly, they typically suppress biased outputs without explicitly modeling the statistical dependence between model outputs and the underlying stereotype concepts. We propose Acmite, a lightweight concept-guided framework for targeted and selective debiasing. Acmite represents stereotypes as structured semantic concepts and uses maximal marginal relevance (MMR) to select diverse concepts for debiasing. Inspired by mutual information minimization, it approximates this dependence with token-level KL divergence while preserving task semantics. A lightweight LoRA adapter is trained with the base model frozen and activated at inference time only when the input is sufficiently similar to stereotype-related concepts; otherwise, the original model is used directly. We evaluate Acmite on BBQ, CrowS-Pairs, and StereoSet, and assess general capability preservation on ARC-Challenge, GSM8K, and PIQA. Experiments across three LLMs show that Acmite effectively mitigates gender bias across complementary evaluation formats while maintaining competitive performance on bias-unrelated tasks. Anonymous code and data are available at https://anonymous.4open.science/r/Acmite-18E2/.
Comments15 pages, 0 figures