arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分布式隐含危害:基于多模态大语言模型(MLLM)的视频 moderation 中的组合安全盲点

Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation

Ruotong Wang, Zihao Zhu, Siwei Lyu, Xin Tao, Baoyuan Wu

arXiv 2609.00206首次发表:更新:

发表机构

The Chinese University of Hong Kong, Shenzhen; State University of New York at Buffalo; Kuaishou Technology(香港中文大学(深圳); 纽约州立大学布法罗分校; 快手科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现基于 MLLM 的视频 moderation 存在分布式隐含危害(DIH)这一组合安全盲点,开发多智能体合成框架构建含超9000个视频的数据集,测试30余种MLLM发现其检测DIH存在显著缺陷。

AI 中文摘要

尽管多模态大语言模型(MLLM)在视频 moderation 中的应用日益广泛,但它们存在一种组合安全盲点:由看似无害的组件构成的视频,作为整体解读时可能传递有害含义。我们将这一现象称为分布式隐含危害(Distributed Implicit Harm, DIH),其中危害源于沿视频分解轴分布的组件之间的关系,而非任何单一明确线索。在众多可能的分解轴中,我们研究了两个代表性案例:视觉片段间的时间分布危害(DIH-T),以及音频与视觉流之间的跨模态危害(DIH-M)。大规模研究和缓解 DIH 需要难以收集的数据:此类视频缺乏组合危害标注,无法通过局部视觉线索、关键词或单模态信号检索,因此不存在于现有安全数据集中。为弥合这一缺口,我们开发了多智能体合成框架,将单独无害的组件组合成有害场景,并生成带有明确推理标注的多样化 DIH 视频,得到了包含超过 9000 个视频的数据集,涵盖纯视觉和视听两种设置。对 30 余种 MLLM(包括前沿专有模型和领先开源系统)的基准测试显示,其在检测 DIH-T 和 DIH-M 方面存在显著且持续的缺陷。值得注意的是,这种缺陷甚至在最强的前沿模型中也存在:它们通常能正确评估孤立的单个组件,但无法识别由其组合产生的有害含义。我们进一步在从社交媒体手动收集的真实世界 DIH 视频集上评估这些模型,观察到了相同的失败模式,凸显 DIH 是视频 moderation 中一个实际且未被充分探索的挑战。

英文摘要

Despite their growing use in video moderation, multimodal large language models (MLLMs) exhibit a compositional safety blind spot: videos composed of seemingly benign components can convey harmful meaning when interpreted as a whole. We refer to this phenomenon as Distributed Implicit Harm (DIH), where harm arises from relations among components distributed along a decomposition axis of the video, rather than from any single explicit cue. Among many possible axes, we study two representative cases: temporally distributed harm across visual segments (DIH-T) and cross-modal harm between audio and visual streams (DIH-M). Studying and mitigating DIH at scale requires data that is difficult to collect: such videos lack compositional harm annotations, evade retrieval based on local visual cues, keywords, or single-modality signals, and are consequently absent from existing safety datasets. To bridge this gap, we develop a multi-agent synthesis framework that composes individually benign components into harmful scenarios and generates diverse DIH videos with explicit reasoning annotations, yielding a dataset of over 9,000 videos spanning visual-only and audio-visual settings. Benchmarking over 30 MLLMs spanning frontier proprietary models and leading open-source systems reveals substantial and consistent deficits in detecting both DIH-T and DIH-M. Notably, this failure persists even among the strongest frontier models: they often correctly assess individual components in isolation but fail to recognize the harmful meaning that emerges from their composition. We further evaluate these models on a manually collected set of real-world DIH videos from social media and observe the same failure mode, highlighting DIH as a practical and underexplored challenge for video moderation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑