语言模型认为谁具备胜任能力?职业偏见的机制分析
Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias
浏览论文内容
中文总结 AI 辅助
本研究通过引入因果框架,揭示语言模型对用户专业知识的内部表征存在职业偏见,该偏见会受人口统计学属性影响,且仅靠行为指标无法完全检测到相关失效模式。
中文摘要 AI 辅助
语言模型(LMs)常通过行为偏见评估,但目前尚不清楚它们是否不再代表导致偏见的潜在关联,还是仅学会了不表达这些偏见。本研究表明,即便行为偏见不可见,表征偏见也常可检测。我们引入一种因果框架,将职业偏见分解为两个测量点:模型对用户胜任能力的内部表征,及其可观测输出。我们推导用户专业知识表征的导向向量,并验证其在问答任务和招聘任务中对模型行为的因果中介作用。将该框架应用于多个开放权重模型后发现,性别、种族、社会经济地位等人口统计学属性会影响模型对用户专业知识的表征,即便行为指标未检测到不同人口群体间的差异。我们还表明,这些模型表征在干预下会影响下游行为,提示仅靠行为指标可能无法检测到的失效模式。
英文摘要
Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them. In this study, we show that representational biases are often detectable, even when behavioral biases are not visible. We introduce a causal framework that decomposes occupational bias into two measurement points: a model's internal representation of a user's competence, and its observable outputs. We derive steering vectors for representations of user expertise, and verify that they causally mediate model behavior in both a question-answering task and a hiring task. Applying this framework to several open-weight models, we find that demographic attributes, such as gender, race, and socioeconomic status, influence a model's representation of user expertise, even in cases where behavioral metrics detect no disparity between demographics. We show that these model representations can influence downstream behavior under intervention, suggesting failure modes that behavioral metrics alone may not detect.
发表机构
- Boston University(波士顿大学)
机构由 AI 辅助整理,请以论文原文为准。