基于大语言模型的主观多偏见检测
Subjective Multi-Bias Detection with Large Language Models
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对文本中的主观多偏见问题,采用大语言模型,在WIKIBIAS数据集上检测三类偏见,相关代码已公开。
AI中文摘要:
本项目研究了文本内容中普遍存在的偏见检测难题,具体聚焦于主观偏见的识别,主观偏见是一种引入不当态度或呈现与实际事实不符的表述的偏见,会损害文本的真实性与可靠性,引发误解和潜在的社会紧张,尤其是当这种偏见通过冒犯性语言表达时。我们遵循前人工作[1],处理文本中三种不同类型的主观偏见:(1)框架偏见,即使用带有特定观点的片面词语或短语;(2)认识论偏见,包含影响文本可信度的微妙语言特征;(3)人口统计学偏见,即基于特定人口因素(如性别或宗教)的预设使用词语或短语。我们使用的输入为可能包含主观偏见的文本,输出为揭示所提供内容中是否存在此类偏见的分类或标注,具体而言,我们在包含4000多个来自维基百科编辑的句子对的语料WIKIBIAS[2]中检测了三种不同类型的多跨度偏见,该数据针对跨度对按偏见类型标注,类别包括:(1)框架偏见,(2)认识论偏见,(3)人口统计学偏见,(4)无偏见。项目代码已在该httpsURL公开。
英文摘要:
In this project, we delved into the pervasive challenge of bias detection within the text content. More specifically, our focus lies on the identification of subjective bias, a type of bias that introduces improper attitudes or portrays a statement at odds with the actual truth. The subjective bias can jeopardize the authenticity and reliability of texts, leading to misconceptions and potential social tensions, especially when expressed through offensive language. Following prior work [1], we tackled with three different types of subjective biases in text: (1) framing bias with the use of one-sided words or phrases containing a particular point of view; (2) epistemological bias which includes subtle linguistic features that can affect the believability of the texts; (3) demographic bias with word/phrase usage under presuppositions of a particular demographic factor (i.e., gender or religion). In terms of the data we utilize, the input consists of texts that may harbor subjective biases. The output is a classification or annotation that reveals the presence or absence of such biases within the provided content. More specifically, we detected three different types of multi-span biases in corpus WIKIBIAS [2] with more than 4,000 sentence pairs from Wikipedia edits. The data is labelled by bias type for span pairs with the following categories: (1) framing bias, (2) epistemological bias, (3) demographic bias, and (4) no bias. The project codes are released at https://github.com/HoningJade/LLM-Bias-Type-Classification.