发表机构
Inner Mongolia University; Xiamen University; University of Sydney; Nanjing University of Science and Technology; Harbin Institute of Technology; Zhejiang University(内蒙古大学; 厦门大学; 悉尼大学; 南京理工大学; 哈尔滨工业大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有语言模型去偏方法的局限,提出HEIMAT启发式自动去偏框架,经实验验证其可在保留NLU性能的同时有效缓解不同文化下的模型偏见。
AI 中文摘要
语言模型(LMs)在预训练过程中常习得各类偏见,并可能在交互中表现出来,进而造成社会危害。现有方法多依赖反事实数据增强或表示投影,但这些策略存在计算成本高、难以扩展到更大模型的实际局限,且多数需要人工标注数据,适用范围局限于特定文化和偏见类别。为克服这些限制,我们提出HEIMAT——一种面向LMs的启发式自动去偏框架,包含两个核心步骤:偏见披露与去偏微调。第一步,它使用简单模板构建启发式提示,用于揭示模型偏见并生成对应上下文提示;第二步,通过最小化这些上下文提示预测的Jensen-Shannon散度来微调模型,以减少偏见。大量实验表明,HEIMAT可在不同文化中有效缓解偏见,同时保持模型的自然语言理解(NLU)性能。
英文摘要
Language models (LMs) often acquire various biases during pre-training and may express them in interactions, potentially causing social harm. Existing methods often rely on counterfactual augmentation or representation projection. These strategies remain limited in practice due to their high computational costs and difficulty in scaling to larger models. Additionally, many of these strategies require manual data annotation, narrowing their scope to specific cultures and bias categories. To overcome these limitations, we propose HEIMAT, a HEurIstic-style autoMATic debiasing framework for LMs. HEIMAT consists of two main steps: bias disclosure and debiasing fine-tuning. In the first step, it uses simple templates to construct heuristic prompts, which are applied to reveal model biases and generate corresponding context prompts. In the second step, it fine-tunes the model by minimizing the Jensen-Shannon divergence of predictions on these context prompts to reduce bias. Extensive experiments show that HEIMAT effectively mitigates bias in different cultures while maintaining the model's natural language understanding (NLU) performance.
Comments13 pages in total, 5 figures