发表机构
Institute for Analytics and Data Science, University of Essex; University of Exeter; Queen Mary University of London(埃塞克斯大学分析与数据科学研究所; 埃克塞特大学; 伦敦大学玛丽皇后学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出CLARA框架,以片段级多模态对齐结合VLM衍生理由,在三个仇恨视频数据集上实现了优于现有最优方法的仇恨视频检测性能。
AI 中文摘要
随着以视频为中心的社交媒体平台快速发展,仇恨言论对个人福祉和社会凝聚力构成严重风险,仇恨视频检测变得愈发重要。与文本或静态多模态内容相比,仇恨视频检测探索不足且更具挑战性,因为仇恨语义常源于语音、音频、视觉等多模态线索的复杂交互,且这些信号往往短暂、隐晦且具有时间依赖性,难以用传统的视频级表示捕捉。本研究提出CLARA,一种用于仇恨视频检测的片段级多模态框架。CLARA不将视频视为单一实例,而是将其建模为细粒度片段序列,能更精准捕捉时间局部化的仇恨信号。我们引入混合专家片段编码器实现自适应多模态对齐,提出局部-全局片段对比目标以联合建模短期线索和长程时间依赖,还通过门控Transformer集成VLM衍生的理由,提供高级语义指导。在三个仇恨视频数据集上的大量实验表明,CLARA始终优于现有最优方法,进一步的 ablation 研究和参数分析验证了各组件的有效性。
英文摘要
Hateful video detection has become increasingly important with the rapid growth of video-centric social media platforms, given the serious risks that hate speech poses to both individual well-being and social cohesion. Compared with text or static multimodal content, hateful video detection remains underexplored and significantly more challenging, as hateful meaning often arises from complex interactions among multimodal cues, including speech, audio, and visual content. Moreover, such signals are often brief, implicit, and temporally dependent, making them difficult to capture using conventional video-level representations. In this work, we propose CLARA, a clip-level multimodal framework for hateful video detection. Instead of treating a video as a single instance, CLARA models it as a sequence of fine-grained clips, enabling more precise capture of temporally localized hateful signals. We introduce a Mixture-of-Experts clip encoder for adaptive multimodal alignment, a local-global segment contrastive objective to jointly model short-term cues and long-range temporal dependencies, and VLM-derived rationales integrated via a gated Transformer to provide high-level semantic guidance. Extensive experiments on three hateful video datasets demonstrate that CLARA consistently outperforms state-of-the-art methods. Further ablation studies and parameter analyses validate the effectiveness of each component.