发表机构
Universidad de Buenos Aires; Buenos Aires AI Safety Hub; The World Bank; Universidad Nacional de General San Martín(布宜诺斯艾利斯大学; 布宜诺斯艾利斯AI安全中心; 世界银行; 国立圣马丁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PowerBench是一个评估语言模型在权力转移请求中偏见的基准,通过区分自我赋权、去权和权力攫取,发现模型拒绝权力攫取多于去权,且存在国籍和用户类型相关的系统性偏见。
AI 中文摘要
语言模型越来越多地协助人们处理与权力相关的请求,因此它们在帮助谁方面的系统性差异可能会在规模上改变权力分配,或者被那些了解到哪些身份被拒绝较少的用户所利用。我们引入了PowerBench,一种针对权力转移请求的评估,它区分了自我赋权、去权(disempowerment)和权力攫取(power grabbing),并包含一个不转移权力的拒绝诱导请求作为对照。我们构建、整理并开源了一个包含此类请求的数据集,这些请求在权力领域、背景、受影响方的规模以及用户先前的权力地位上有所变化,并在三种实验条件下评估了24个模型(12个来自美国开发者,12个来自中国开发者):用户与受影响方的国籍互相对应、用户为AI代理,以及8种请求语言。模型对权力攫取的拒绝多于去权,对去权的拒绝多于自我赋权。对权力攫取的拒绝随受影响方规模(从个人到社会)的增加而上升。模型倾向于帮助他人从美国夺取权力,而不愿帮助美国用户从他人那里夺取权力,但当美国获得权力且无人失去时,模型则偏向美国。当用户是AI代理时,对权力转移请求的拒绝增加,尤其是在针对个人的权力攫取中。最后,语言影响拒绝行为,但以模型特定的方式,平均而言在很大程度上相互抵消。我们发布PowerBench,以使这些不对称性在当前和未来的模型中可衡量。
英文摘要
Language models increasingly assist people with power-related requests, so systematic differences in whom they help could shift the distribution of power at scale, or be exploited by users who learn which identities are refused less. We introduce PowerBench, an evaluation of power-shifting requests that distinguishes self-empowerment, disempowerment, and power grabbing, plus a control of refusal-inducing requests that shift no power. We build, curate, and open-source a dataset of such requests varying the power domain, the context, the scale of the affected party, and the prior power standing of the user, and evaluate 24 models (12 from US and 12 from Chinese developers) under three experimental conditions: reciprocal nationalities of user and affected party, an AI agent as the user, and 8 request languages. Models refuse power grabbing more than disempowerment, and disempowerment more than self-empowerment. Refusal of power grabbing rises with the scale of the affected party, from an individual to a society. Models are biased toward helping others take power from the US and against helping US users take power from others, but favor the US when it gains power and nobody loses it. When the user is an AI agent, refusal of power-shifting requests increases, especially in power grabbing against an individual. Finally, language biases refusal, but in model-specific ways that largely cancel on average. We release PowerBench to make these asymmetries measurable in current and future models.