Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying Prompts
双向偏见归因:无需修改提示的大型语言模型去偏方法
机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) ; vivo AI Lab, China(vivo AI实验室) ; Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室(厦门大学),文化部,中国)
专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI
AI总结 本文提出一种无需微调或修改提示的双向偏见归因方法,通过检测刻板印象词汇并归因神经元偏见,有效减少LLM中的偏见同时保持性能。