arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一个示例足以通过公平性基准:重新思考对齐大语言模型的公平性评估

One Example Is Enough to Pass Fairness Benchmarks: Rethinking Fairness Evaluation for Aligned LLMs

Naihao Deng, Samee Arif, Shuaichen Chang, Yulong Chen, Rada Mihalcea

arXiv 2609.14860首次发表:更新:

发表机构

University of Michigan; The Ohio State University; University of Aberdeen; University of Cambridge(密歇根大学; 俄亥俄州立大学; 阿伯丁大学; 剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过实验证明单个示例即可大幅提升BBQ公平性基准得分,指出该类基准仅衡量单一结构线索,无法真实反映模型公平性,呼吁构建更全面的评估套件。

AI 中文摘要

警告:本研究涉及刻板印象和偏见,包含有毒和冒犯性示例,仅用于说明目的。诸如BBQ之类的公平性基准已成为各大模型家族公平性评估的事实标准。我们认为这些基准过于简单,无法支撑其作用:使用组相对策略优化(GRPO)在单个BBQ示例上训练Qwen 2.5 7B Base,或将此示例作为上下文中的一次性演示用于上下文学习(ICL),分别将平均BBQ准确率从79.9%提升至92.9%和99.0%,通过GRPO缩小了与其大规模RLHF对应模型(96.1%)差距的80%,并通过ICL超越了该模型。这些效应在多个模型家族中普遍存在。交叉条件分析表明,改进由模型生成的推理轨迹所承载,一个示例足以引发一种与类别无关的“缺失证据”推理模式。我们认为,BBQ式的多项选择弃权(不执行)基准仅衡量单一结构线索,解决这些基准的模型并不因此变得公平。我们呼吁建立覆盖更广泛公平性对齐维度的评估套件。

英文摘要

Warning: This submission studies stereotypes and biases, and contains toxic and offensive examples, used for illustration purposes only. Fairness benchmarks such as BBQ have become the de facto standard for fairness evaluation across major model families. We argue that these benchmarks are too easy to support their role: training Qwen 2.5 7B Base with Group Relative Policy Optimization (GRPO) on a single BBQ example, or placing that example in context as a one-shot demonstration for in-context learning (ICL), lifts mean BBQ accuracy from 79.9% to 92.9% and 99.0%, respectively, closing 80% of the gap to its large-scale RLHF counterpart (96.1%) with GRPO, and surpassing it with ICL. These effects generalize across model families. A cross-conditioning analysis shows the improvement is carried by the reasoning traces generated by the model, and one example suffices to elicit a category-agnostic ``missing evidence'' reasoning pattern. We argue that BBQ-style multiple-choice abstention benchmarks measure a single structural cue, and a model that solves them does not thereby become fair. We call for evaluation suites that cover a broader spectrum of fairness alignment.

CommentsAccepted to EMNLP 2026 Main

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑