发表机构
MPI-SWS; Saarland University; INRIA(马克斯·普朗克软件系统研究所; 萨尔大学; 法国国家信息与自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过新基准ModerationBench比较指令驱动与示例驱动范式,证明基础模型在内容审核中显著优于现有系统,F1分数近三倍提升,为可靠政策执行提供路径。
AI 中文摘要
内容审核政策日益复杂,对其一致执行构成了严峻挑战。尽管基础模型具备应对这一挑战的基本能力,但它们能否可靠地审核在线内容仍是一个悬而未决的问题。在本文中,我们系统比较了视觉-语言模型(VLM)指导的两种竞争范式:指令驱动方法(模型依据政策准则进行推理)与示例驱动方法(模型从先前先例中泛化)。我们将这一研究建立在ModerationBench之上,这是一个包含来自Bluesky平台的4,000条人工标注、真实世界帖子的新基准。我们的实验表明,基础模型能大幅超越Bluesky部署的审核系统,在基准的随机帖子上将其F1分数几乎提高三倍(0.60对0.22),且指令驱动与示例驱动范式均达到了相当的峰值效果。因此,我们的发现为大规模可靠且适应性强的政策执行指明了一条路径。
英文摘要
The growing complexity of content moderation policies presents a critical challenge for their consistent operationalization. While foundation models possess the basic capabilities needed to confront this challenge, whether they can reliably moderate online content remains an unanswered question. In this paper, we systematically compare two competing paradigms for Vision-Language Model (VLM) guidance: an instruction-driven approach where models reason from policy precepts, and an example-driven approach where they generalize from prior precedents. We ground this investigation in ModerationBench, a new benchmark of 4,000 manually annotated, in-the-wild posts from the Bluesky platform. Our experiments reveal that foundation models can substantially outperform Bluesky's deployed moderation system, nearly tripling its $F_1$ score (0.60 vs. 0.22) on Random Posts in the benchmark, with both instruction- and example-driven paradigms achieving comparable peak effectiveness. Our findings thus chart a path toward reliable and adaptable policy operationalization at scale.
Comments33 pages, 28 figures, 8 tables