arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40124cs.CL

自行去偏:教大语言模型认知偏差缓解干预措施

Debias It Yourself: Teaching LLMs Cognitive Bias Mitigation Interventions

  • George Mason University(乔治梅森大学)
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

Chahat Raj, Sina Mansouri, Aylin Caliskan, Antonios Anastasopoulos, Ziwei Zhu

AI总结:

提出基于认知的DIY框架,将五种人类去偏干预转化为大语言模型的去偏程序,通过展示、训练、修订三种范式实施,在多个基准上显著降低偏见并保持推理性能。

AI中文摘要:

偏见在社会心理学和认知科学中已被长期研究,数十年的研究产生了一系列经过验证的干预措施,这些措施能减少人类的刻板思维和偏见反应。我们提出“自行去偏”(DIY),一个基于认知的框架,将五种此类干预措施转化为大语言模型的去偏程序,并通过三种既定范式实施:展示(上下文示例)、训练(指令微调)和修订(引导式自我修订)。在三个模型、五个偏见基准、十一个去偏基线和三个推理基准上,训练+修订和仅修订取得了前两名的平均排名,在偏见-推理权衡中领先(在90%推理准确率下平均偏见低至2%),并在未见维度上将偏见最多降低14.8%。我们的代码和数据公开可用。

英文摘要:

Bias has long been studied in social psychology and cognitive science, where decades of research have produced a body of validated interventions that reduce stereotypical thinking and prejudiced responses in humans. We propose Debias It Yourself (DIY), a cognitively grounded framework that translates five such interventions into debiasing procedures for large language models and delivers them through three established paradigms: Show (in-context examples), Train (instruction tuning), and Revise (guided self-revision). Across three models, five bias benchmarks, eleven debiasing baselines, and three reasoning benchmarks, Train+Revise and Revise alone attain the top two average ranks, lead the bias-reasoning tradeoff (mean bias as low as 2% at 90% reasoning accuracy), and reduce bias on unseen dimensions by up to 14.8%. Our code and data are publicly available.

补充信息

↑