arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分析与修正大语言模型中的善意偏差

Analyzing and Correcting Benevolence Bias in Large Language Models

Yuanzi Li, Junhao Wang, Minghui Liu, Boyi Li, Bingchen Chen, Zihang Tian, Jingyu Zhao, Yuhan Wang, Lei Wang, Pei Wang, Jinchao Wu, Xu Chen

arXiv 2608.24912首次发表:更新:

发表机构

Gaoling School of Artificial Intelligence, Renmin University of China; School of Computer Science and Engineering, Sun Yat-sen University; School of Computer Science and Technology, Shandong University; School of Mathematical Sciences, Peking University(中国人民大学高瓴人工智能学院; 中山大学计算机科学与工程学院; 山东大学计算机科学与技术学院; 北京大学数学科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究识别出大语言模型存在的善意偏差,发现其稳定且随模型规模增大而增强,提出轻量对比校准法可修正该偏差,明确了对齐后的大语言模型作为人类替代者的适用场景与修正方法。

AI 中文摘要

大语言模型(LLM)越来越多地被用作人类受访者的替代者,应用于民意调查、模拟调查参与者以及基于智能体的社会模拟等场景。这些应用基于一个假设:让模型适配某类人群的身份,其生成的答案会与该群体真实人类的答案相似。本文识别并测量了善意偏差,这是对齐后的LLM在涉及价值观的调查问题上,倾向于选择更友善、更安全、更符合社会认可的答案的一种虽小但稳定的趋势。在18种广泛使用的模型、4个社会科学数据集(ANES、GSS、WVS及一个跨文化前景理论复现数据集)和6个心理学类别上,我们发现该偏差是模型的稳定属性,而非单一系统的怪癖:它在不同模型中方向一致,随模型规模增大而增强,且源于后训练阶段。提示语言和框架会改变偏差的大小,但永远不会改变其方向;“恶意角色”压力测试显示出单向限制:对齐后的模型难以扮演比平均水平更不友善、更亲社会程度更低或更能容忍伤害的角色。因此,问题不仅是平均值偏移,还包括模型可模仿的人群范围缩小。该偏差位于答案分布的中间而非尾部,且在采样温度变化和简单提示反思后仍存在。令人鼓舞的是,它易于诊断且修正方法直接:轻量对比校准无需重新训练,可应用于黑盒API,能将全部6个类别拉回人类基准线。我们的结果为研究人员提供了清晰的图谱,明确对齐后的LLM可在哪些场景作为人类替代者被信任、哪些场景需要注意,并提供了缩小差距的现成可用方法。

英文摘要

Large language models (LLMs) are increasingly used as stand-ins for human respondents, from opinion polls and simulated survey participants to agent-based social simulations. These uses rest on one assumption: that conditioning a model on who a person is yields answers resembling those of real people from that group. Here we identify and measure benevolence bias, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions. Across 18 widely used models, four social-science datasets (ANES, GSS, WVS, and a cross-cultural prospect-theory replication) and six psychological categories, we find that the bias is a stable model property, not a quirk of any one system: it points the same way across models, grows with model size, and traces to the post-training stage. Prompt language and framing change its size but never its direction, and a "malicious persona" stress test shows a one-sided limit: aligned models struggle to play people who are less kind, less prosocial or more harm-tolerant than average. The issue is thus not only a shifted average, but a narrowed range of people the model can imitate. The bias sits in the middle of the answer distribution rather than its tails, and survives changes in sampling temperature and simple prompted reflection. The encouraging news is that it is easy to diagnose and straightforward to fix: a light-touch contrastive calibration, which needs no retraining and works on black-box APIs, brings all six categories back to the human baseline. Our results give researchers a clear map of where aligned LLMs can already be trusted as human stand-ins, where they need care, and a ready-to-use method for closing the gap.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑