AI 中文总结
该研究通过黑盒评估发现,商用多模态图像审核API可被简单图像变换绕过,且不同服务稳健性差异大,表明这类API无法单独作为可靠安全过滤器,需结合分层审核管道部署。
AI 中文摘要
自动化内容审核系统已成为大规模筛查有害内容的关键,然而传统的任务特定分类器往往提供有限的政策覆盖范围和上下文理解能力。近来,基于大型基础模型构建的商用多模态审核API被推出,承诺提供更广泛、更强大的安全过滤器。在本研究中,我们分析这一转变是否也带来了更稳健的图像审核效果。我们对三个已建立的商用图像审核服务开展大规模黑盒评估,并比较它们的稳健性。通过在多个提供商、数据集、伤害类别、感知相似性约束和变换强度下评估七种简单的、与模型无关的图像变换,我们发现:(1)所有三个商用服务都可通过无需梯度、替代模型或目标系统知识的低成本图像变换被绕过;(2)即使是颜色反转和灰度转换等固定变换,也会引发不安全到安全的决策变化,同时保留人类仍可识别的内容;(3)它们在不同数据集和伤害类别间的稳健性差异显著,其中多模态内容和自伤表现出明显的漏洞。由此得出结论:用基于基础模型的API替代传统审核分类器本身并不能提供可靠的安全边界。这类系统必须在现实变换下进行评估,并作为分层审核管道的一个组件部署,而非独立的安全过滤器。
英文摘要
While automated content-moderation systems have become essential for screening harmful content at scale, conventional task-specific classifiers often provide limited policy cov- erage and contextual understanding. Recently, commercial multimodal moderation APIs built on large foundation models have been introduced with the promise of providing broader and more capable safety filters. In this work, we analyze whether this shift also yields more robust image moderation. We conduct a large-scale black-box evaluation on three established commercial image-moderation services and compare their robustness. By evaluating seven simple, model-agnostic image transformations across multiple providers, datasets, harm categories, perceptual-similarity constraints, and transformation intensities, we find that: (1) all three commercial services can be bypassed using inexpensive image transformations that require no gradients, surrogate models, or knowledge of the target system; (2) even fixed transformations such as color inversion and grayscale conversion induce unsafe-to-safe decision changes while preserving content that remains recognizable to humans; (3) their robustness varies substantially across datasets and harm categories, with multimodal content and self-harm exhibiting pronounced vulnerabilities. This yields the conclusion that replacing conventional moderation classifiers with foundation-model-based APIs does not, by itself, provide a reliable security boundary. Such systems must be evaluated under realistic transformations and deployed as one component of a layered moderation pipeline rather than as standalone safety filters.