发表机构
University of Vermont; Santa Fe Institute; Complexity Science Hub(佛蒙特大学; 圣塔菲研究所; 复杂性科学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究借助边缘归因修补(EAP)定位了GPT-2 Small等模型中“我们对他们”偏见的有向回路,发现该偏见与语言能力通路纠缠,且不同模型的内部组织存在差异。
AI 中文摘要
大型语言模型(LLMs)存在“我们对他们”偏见:一种行为不对称性,即以内群体(“我们”/“us”)为框架的提示,相比以匹配的外群体(“他们”/“them”)为框架的提示,会系统地获得更积极的续文。使用边缘归因修补(EAP),我们将此行为定位在GPT-2 Small、GPT-2 Large、Llama-2 7B和Llama-3 8B的有向回路中。在所有模型中,高归因边缘稀疏且主要集中在中后层,而与内群体和外群体相关的通路仅部分重叠,且在微调下会重组。行为比较和30-token生成结果表明,测得的不对称性超出了下一个token评估的范围。对EAP排名边缘的针对性修补,相比等基数的随机边缘对照,能更大程度降低偏见分数,但也会显著降低CoLA、HellaSwag和WikiText-2的性能,这表明与偏见相关的通路与更广泛的语言能力在功能上相互纠缠。总体而言,我们的研究结果将“我们对他们”偏见描述为一种稀疏、依赖上下文且功能纠缠的计算现象,其内部组织在所检查的模型中存在差异。
英文摘要
Large language models (LLMs) exhibit us-vs.-them bias: A behavioral asymmetry in which prompts framed around an ingroup (``we''/``us'') receive systematically more positive continuations than matched prompts framed around an outgroup (``they''/``them''). Using Edge Attribution Patching (EAP), we localize this behavior to directed circuits in GPT-2 Small, GPT-2 Large, Llama-2 7B, and Llama-3 8B. Across models, high-attribution edges are sparse and concentrated primarily in middle-to-late layers, while ingroup- and outgroup-related pathways are only partially overlapping and reorganize under fine-tuning. Behavioral comparisons and 30-token generations show that the measured asymmetry extends beyond next-token evaluation. Targeted patching of EAP-ranked edges reduces the bias score more than equal-cardinality random-edge controls, but also substantially degrades performance on CoLA, HellaSwag, and WikiText-2, suggesting that bias-relevant pathways are functionally entangled with broader language capabilities. Overall, our findings characterize us-vs.-them bias as a sparse, context-dependent, and functionally entangled computational phenomenon whose internal organization varies across the models examined.
Comments13 page main paper with 3 figures, 2 tables; 15 page appendix with 9 figures, 7 tables