发表机构
Utrecht University; Wageningen University & Research(乌得勒支大学; 瓦赫宁根大学及研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对性别歧视检测中标注分歧问题,提出MAP-PO框架,通过聚类标注者并训练对应智能体,结合个体与团队奖励协调智能体,实验验证了聚类训练及团队信号的必要性。
AI 中文摘要
当人们对文本进行性别歧视标注时,往往会存在分歧,这并非因为部分人判断错误,而是因为他们对性别歧视的感知确实存在差异。大多数自然语言处理(NLP)系统会通过将分歧简化为多数投票来忽略这些差异。我们提出了多智能体视角偏好优化(MAP-PO)框架,以保留这些不同的视角。在带有标注的英语和西班牙语推文的EXIST 2024数据集上,我们首先根据标注行为而非人口统计学属性对标注者进行聚类。随后,我们为每个聚类微调一个大语言模型智能体,以复现该聚类的标注行为,并结合个体和团队层面的奖励通过偏好优化来协调这些智能体。我们在由两种语言和两种主干语言模型定义的四种设置中评估MAP-PO,以探究每个智能体是否能复现其所属聚类的标注结果,以及这些智能体整体是否能复现多数标注结果。在所有四种设置中均得出两个结论:第一,未经过微调时,智能体的行为几乎完全相同,因此聚类特定训练是必要的;第二,我们发现仅在自身聚类的标注上训练每个智能体,会使智能体偏离其应代表的聚类,而添加共享的团队层面训练信号则能持续让每个智能体与所属聚类保持校准。
英文摘要
When people label text for sexism, they often disagree, and not because some of them are wrong: they genuinely perceive sexism differently. Most NLP systems discard this disagreement by collapsing it into a majority vote. We propose the Multi-Agent Perspectivist Preference Optimization (MAP-PO) framework to keep these different perspectives. On the EXIST 2024 dataset of labeled English and Spanish tweets, we first cluster annotators by their labeling behavior rather than their demographic attributes. We then fine-tune one Large Language Model agent per cluster to reproduce that cluster's annotation behavior, and coordinate the agents with preference optimization that combines individual and team-level rewards. We evaluate MAP-PO in four settings defined by two languages and two backbone language models, asking whether each agent reproduces the annotations of its own cluster and whether the agents together reproduce the majority label. Two findings hold in all four settings. First, without fine-tuning the agents behave almost identically, so cluster-specific training is necessary. Second, we show that training each agent only on the labels of its own cluster pushes the agents far beyond the clusters they should represent, while adding a shared team-level training signal consistently keeps each agent calibrated to its cluster.
Comments17 pages, 12 figures, 14 tables. Preprint; under review at EACL 2027 (ACL Rolling Review, August 2026 cycle). Code and data: https://github.com/mohammadi-hadi/MAP-PO