谁的判断算数?众包内容审核中的代表性差距导致对感知毒性的保护不平等
Whose Judgments Count? Representation Gaps in Crowdsourced Content Moderation Produce Unequal Protection from Perceived Toxicity
浏览论文内容
中文总结 AI 辅助
该研究结合大规模判断数据与反事实模拟,发现众包内容审核存在同群体保护模式,审核员人口构成会导致对感知毒性的保护不平等,黑人及LGB群体需远超其人口占比的代表才能获平等保护。
中文摘要 AI 辅助
内容审核是数字治理的核心形式,但人们对于应从共享网络空间移除哪些内容存在分歧。各平台虽会汇总人类判断来构建审核系统,然而这一过程如何影响哪些用户能免受其感知为有毒的内容侵害,目前仍不明确。我们结合大规模判断数据与反事实模拟,探究审核员群体的人口统计构成如何塑造对用户的保护分布,以此填补这一研究空白。将该框架应用于16221名美国受访者对Twitter、Reddit及4chan上102463条评论的移除判断后,我们发现审核需求存在人口统计异质性。我们进一步揭示了一致的同群体保护模式:感知毒性的减少会不成比例地惠及与审核员群体拥有相同人口统计身份的用户。至关重要的是,与全国代表性基准相比,反映Prolific平台自审员人口统计构成的审核员群体会扩大这些差距,而即便是完全代表性的群体也无法确保平等保护:除非黑人及LGB群体的代表性远超其人口占比,否则他们仍会受到保护不足。这些发现表明,对感知毒性的不平等保护可从结构上源于分层移除标准的汇总,使审核输入的人口统计构成成为决定谁能在线获得保护的关键因素。
英文摘要
Content moderation is a central form of digital governance, yet people disagree over what content should be removed from shared online spaces. While platforms aggregate human judgments to build moderation systems, it remains unclear how this process shapes which users are protected from content they perceive as toxic. We address this gap by combining large-scale judgment data with counterfactual simulations that trace how the demographic composition of moderator pools shapes the distribution of protection across users. Applying this framework to removal judgments from 16,221 U.S. respondents evaluating 102,463 comments from Twitter, Reddit, and 4chan, we find demographic heterogeneities in moderation demand. We further reveal a consistent pattern of in-group protection: reductions in perceived toxicity accrue disproportionately to users who share the demographic identities of the moderator pool. Crucially, moderator pools that mirror the demographic composition of self-identified moderators on Prolific widen these disparities relative to a nationally representative baseline, while even fully representative pools fail to ensure equal protection: Black and LGB users remain underprotected unless they are represented well beyond their population share. These findings show that unequal protection from perceived toxicity can arise structurally from the aggregation of stratified removal standards, making the demographic composition of moderation inputs a key determinant of who is protected online.
发表机构
- University of South Carolina(南卡罗来纳大学)
- New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。