arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2610.08747cs.CL

缺失的最小对:大语言模型中的刻板印象评估

The Missing Minimal Pair: Stereotype Evaluation in LLMs

  • University of Edinburgh(爱丁堡大学)
  • University of Amsterdam(阿姆斯特丹大学)

机构由 AI 辅助整理,请以论文原文为准。

Nataliya Stepanova, Ivan Titov, Emily Allaway, Björn Ross

AI总结:

针对大语言模型刻板印象评估中单对比较不可靠的问题,提出双最小对设置及数据增强框架,并引入基于互信息的评估指标,实现跨语言和模型的稳健比较。

AI中文摘要:

衡量大语言模型偏见的一种常见方法是比较两个对比性刻板印象句子的对数似然。我们认为,这种单对比较往往不可靠:仅仅用另一个属性改写同一个刻板印象,就可能产生逻辑上不一致的偏好。为解决这一问题,我们提出了一种双最小对设置,引入两个比较轴以实现稳健的刻板印象评估。首先,我们提出了一个数据增强框架,通过生成改写和替代属性来填补现有刻板印象数据集中的关键空白。我们将该框架应用于一组英语、俄语、西班牙语和汉语刻板印象。其次,我们引入了两种针对双最小对设置量身定制的评估指标。其中一种指标通过建模社会群体与刻板化属性之间的互信息(MI),为偏见提供了新的视角。这种基于MI的指标更适合聚合,并能实现跨语言和跨模型更稳健的刻板印象强度比较。我们的代码可在以下网址获取:https://this https URL。

英文摘要:

A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two contrastive stereotype sentences. We argue that such single-pair comparisons are often unreliable: simply rewriting the same stereotype with an alternative attribute can yield logically inconsistent preferences. To address this, we propose a dual minimal pair setup that introduces two axes of comparison for robust stereotype evaluation. First, we present a data-augmentation framework that fills critical gaps in existing stereotype datasets by generating paraphrases and alternate attributes. We apply our framework on a set of English, Russian, Spanish and Chinese stereotypes. Second, we introduce two evaluation metrics tailored to the dual minimal pair setup. One of these metrics provides a new perspective on bias by modeling the mutual information (MI) between social groups and stereotyped attributes. This MI-based metric is better suited for aggregation and enables more robust comparisons of stereotype strength across different languages and models. Our code is available at https://github.com/stepanat/missing-minimal-pair/.

↑