探究基于推理的语言模型在缓解社会偏见中的思维行为
Investigating Thinking Behaviours of Reasoning-Based Language Models for Social Bias Mitigation
AI总结:
本文研究了基于推理的语言模型在缓解社会偏见中的思维机制,发现两种导致偏见累积的失败模式,并提出轻量级提示方法减少偏见同时保持准确性。
AI中文摘要:
尽管基于推理的大语言模型通过内部结构化思维过程在复杂任务中表现出色,但这种思维过程可能聚合社会刻板印象,导致偏见结果。然而,这些语言模型在社会偏见场景中的底层行为仍缺乏研究。本文系统地探究了这一现象背后的机制,揭示了两种驱动社会偏见聚合的失败模式:1) 刻板印象重复,即模型主要依赖社会刻板印象作为主要依据;2) 无关信息注入,即模型伪造或引入新细节以支持偏见叙述。基于这些见解,我们提出了一种轻量级提示方法,通过查询模型对其初始推理进行审查,以应对这些特定的失败模式。在问答(BBQ和StereoSet)和开放性(BOLD)基准测试中,我们的方法有效减少了偏见,同时保持或提高了准确性。
英文摘要:
While reasoning-based large language models excel at complex tasks through an internal, structured thinking process, a concerning phenomenon has emerged that such a thinking process can aggregate social stereotypes, leading to biased outcomes. However, the underlying behaviours of these language models in social bias scenarios remain underexplored. In this work, we systematically investigate mechanisms within the thinking process behind this phenomenon and uncover two failure patterns that drive social bias aggregation: 1) stereotype repetition, where the model relies on social stereotypes as its primary justification, and 2) irrelevant information injection, where it fabricates or introduces new details to support a biased narrative. Building on these insights, we introduce a lightweight prompt-based mitigation approach that queries the model to review its own initial reasoning against these specific failure patterns. Experiments on question answering (BBQ and StereoSet) and open-ended (BOLD) benchmarks show that our approach effectively reduces bias while maintaining or improving accuracy.