arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

衡量LLM中的人类类似偏见?对LLM评估中源自人类偏见构念的批判

Measuring Human-Like Bias in LLMs? A Critique of Human-Derived Bias Constructs in LLM Evaluation

Antonela Tommasel, Markus Schedl

arXiv 2610.00070首次发表:更新:

发表机构

Johannes Kepler University Linz; ISISTAN, CONICET-UNICEN; Linz Institute of Technology(林茨约翰开普勒大学; ISISTAN,CONICET-UNICEN; 林茨理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文批判了用人类偏见构念评估LLM的推断差距,提出一个框架以明确目标构念、操作化、推断范围及人类类比局限,从而合理阐释模型偏见。

AI 中文摘要

研究者越来越多地使用源自人类的偏见构念来研究大语言模型(LLMs),包括社会认知构念如内隐偏见和刻板印象激活,以及认知偏见如锚定效应、框架效应和确认偏见。此类方法为直接的偏见探测提供了替代方案,尤其是在直接提问可能掩盖偏见或模型行为看似规范可接受的情况下。然而,将人类偏见构念适配到LLMs引入了一个推断差距。心理测量工具是为研究人类认知和社会行为而开发的,而LLM评估依赖于概率、文本补全、排序或模拟决策。本文批判了LLM中以人为中心的偏见评估。我们展示了这一差距如何源于与人类衍生构念、人类-模型差异以及评估情境相关的不匹配,这些不匹配可能模糊对模型偏见的不同解释。随后,我们引入了一个框架,提供一个分析视角,将这些要素与合理的解释联系起来,关注目标构念、操作化、推断范围以及人类类比的局限性。

英文摘要

Researchers increasingly use human-derived bias constructs to study Large Language Models (LLMs), including social-cognitive constructs such as implicit bias and stereotype activation, and cognitive biases such as anchoring, framing effects, and confirmation bias. Such approaches offer alternatives to overt bias probes, particularly when direct questioning may obscure bias or when model behaviour appears normatively acceptable. However, adapting human bias constructs to LLMs introduces an inferential gap. Psychological instruments were developed to study human cognition and social behaviour, whereas LLM evaluations rely on probabilities, text completions, rankings, or simulated decisions. This paper critiques human-centered bias evaluation in LLMs. We show how this gap arises from mismatches pertaining to human-derived constructs, human-model differences, and evaluation contexts, which can blur distinct interpretations of model bias. We then introduce a framework providing an analytical lens for relating these elements to warranted interpretations, with attention to target constructs, operationalizations, scope of inference, and limits of human analogy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑