刻画 Bluesky 内容审核服务:从服务的自动化到危害的图景
Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms
- Max Planck Institute for Software Systems(马克斯·普朗克软件系统研究所)
- Saarland University(萨尔大学)
- Indian Institute of Technology Kharagpur(印度理工学院卡拉格普尔分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究利用 Bluesky 的透明架构,首次大规模审计其默认审核服务,分析 1060 万条标签,揭示其高精确率低召回率的人机协作机制及危害图景,为设计更优审核系统提供数据基础。
AI中文摘要:
对内容审核的实证研究从根本上受到主流社交媒体平台上审核系统部署不透明的制约。为此,近期出现的具有透明、公开审核日志的去中心化平台为独立审计提供了前所未有的机会。在本工作中,我们利用这种架构透明度,对 Bluesky 上的默认审核系统——Bluesky 审核服务(BMS)——进行了首次大规模审计。通过分析其 2025 年的 1060 万条审核标签,我们研究了三个基本方面:(i)其机制(自动化与人工监督的程度);(ii)其效能(检测危害的准确性);以及(iii)其目的(其识别出的危害图景)。我们的研究结果揭示了一个人机协作系统,其中针对色情和露骨内容的标签在数秒内自动应用,而细微且高风险的标签则需要更多人工监督,耗时数小时或数天。通过一项人工标注研究,我们发现 BMS 以高精确率(0.837)运行,但召回率较低(0.222),我们的标注员在随机样本中识别出的有害内容比审核系统多 4.5 倍。最后,对最频繁应用的已标注帖子的无监督聚类揭示了检测到的危害,范围从针对受保护群体的敌意言论到色情露骨内容及其他露骨内容的传播。我们的工作深入了解了已部署审核系统的运营现实,为设计更有效、更透明的审核系统提供了具体的数据驱动基础。
英文摘要:
Empirical research on content moderation is fundamentally constrained by the opaque deployment of moderation systems on major social media platforms. To this end, the recent emergence of decentralized platforms with transparent, public moderation logs presents an unprecedented opportunity for independent audits. In this work, we leverage this architectural transparency to conduct the first large-scale audit of the default moderation system on Bluesky, the Bluesky Moderation Service (BMS). Analyzing its 10.6M moderation labels from 2025, we investigate three foundational aspects: (i) its mechanism (the degree of automation versus human oversight), (ii) its efficacy (accuracy in detecting harms), and (iii) its purpose (the landscape of harms it identifies). Our findings reveal a human-AI collaborative system where labels for sexual and graphic content are applied automatically in seconds, while nuanced and high stakes labels require more human oversight, taking hours or days. Through a manual annotation study, we find the BMS operates with high precision (0.837), but struggles with low recall (0.222), with our annotators identifying 4.5$\times$ more harmful content than the moderation system in a random sample. Finally, unsupervised clustering of the most frequently applied labeled posts uncovers detected harms ranging from hostility in discourse toward protected groups to the spread of sexually explicit and other graphic content. Our work offers a look into the operational realities of a deployed moderation system, providing a concrete data-driven foundation for designing more effective and transparent moderation systems.