感知可靠性的性别歧视检测:将DPO与标注者一致性及词元级置信度评分相结合
Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring
浏览论文内容
中文总结 AI 辅助
该研究针对在线性别歧视检测的主观性问题,提出RA-DPO方法,结合标注者一致性等信号,在EXIST 2023数据集上验证其可降低训练成本并提升推理准确率。
中文摘要 AI 辅助
在线性别歧视检测仍是一个未解决的问题。性别歧视检测本质上具有主观性,但现有大多数系统将多个标注者的标签简化为单一多数决策,并对所有实例一视同仁,忽略了两个有价值的信号:标注者一致性和模型不确定性。我们提出RA-DPO(感知可靠性的直接偏好优化),该方法将标注者一致性、模型置信度和词元级不确定性信号整合为单一可靠性评分。RA-DPO利用该评分在训练期间选择高价值偏好对,并支持推理时弃权(不执行),使模型能够在覆盖范围和准确性之间进行权衡。我们在来自EXIST 2023的6920条多语言帖子上对RA-DPO进行评估,通过DPO微调OpenAI gpt-4o base,并在两个开源权重的3B模型(Llama、Qwen)上进行验证。结果显示,基于前30%最可靠的偏好对进行训练的效果与使用全部数据进行DPO训练的效果相当,这表明感知可靠性的选择可以在不牺牲性能的情况下降低训练成本。在推理阶段,选择性预测在真实一致性设置下,当覆盖范围为50%时准确率达到96.2%;在可部署的预测一致性设置下,准确率为88.7%,均超过了85.3%的无一致性基线。这些结果表明,考虑标注不确定性对于主观分类的高效训练和可靠部署均有益处。
英文摘要
The detection of online sexism remains an open problem. Sexism detection is inherently subjective, yet most existing systems reduce multi-annotator labels to a single majority decision and treat all instances uniformly. This ignores two informative signals: annotator agreement and model uncertainty. We propose RA-DPO (Reliability-Aware Direct Preference Optimization), which integrates annotator agreement, model confidence, and a token-level uncertainty signal into a single reliability score. RA-DPO uses this score to select high-value preference pairs during training and to support inference-time abstention, which allows the model to trade coverage for accuracy. We evaluate RA-DPO on 6,920 multilingual posts from EXIST 2023, fine-tune OpenAI gpt-4o base via DPO, and validate on two open-weight 3B models (Llama, Qwen). Results show that training on the top 30% most reliable pairs matches full-data DPO, which indicates that reliability-aware selection can reduce training cost without sacrificing performance. At inference, selective prediction reaches 96.2% accuracy at 50% coverage in the true-agreement setting and 88.7% in the deployable predicted-agreement setting, both exceeding the 85.3% no-agreement baseline. These results suggest that accounting for annotation uncertainty is beneficial for both efficient training and reliable deployment in subjective classification.
发表机构
- Utrecht University(乌得勒支大学)
- Wageningen University & Research(瓦赫宁根大学及研究中心)
机构由 AI 辅助整理,请以论文原文为准。