arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02019cs.CLcs.CV

基于自适应Tversky策略优化的可控多标签视频安全检测

Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization

Guangyu Yang, Jingbiao Mei, Mingsheng Sun, Jinghong Chen, Yingtong Bu, Pengda Qin, Da Chen, Bill Byrne

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有视频安全检测忽视多标签特性和静态训练目标的问题,提出自适应Tversky策略优化(ATPO)框架,通过动态调整假阳性和假阴性惩罚实现可控精确率-召回率权衡,在SafeWatch-Bench上显著提升多标签性能。

中文摘要 AI 辅助

视频社交媒体的快速增长增加了用户接触有害内容的风险,因此需要可靠的自动化视频安全检测系统。尽管近期视觉语言模型(VLMs)展现出强大的视频理解能力,但现有的有害视频检测系统存在两个关键局限:它们通常将安全检测简化为二分类问题,忽视了不安全视频固有的多标签特性;同时,它们依赖静态训练目标,无法支持可控的精确率-召回率权衡,而不同审核流程和不安全类别可能对工作点有不同的需求。为解决这些问题,我们提出了自适应Tversky策略优化(ATPO),一种用于多标签视频安全检测(Multi-VSD)的强化学习框架。ATPO引入了自适应Tversky奖励(ATR),在训练过程中动态调整假阳性和假阴性惩罚,以实现可控的精确率-召回率权衡。在SafeWatch-Bench和XD-Violence上的实验表明,ATPO显著提升了多标签性能,在SafeWatch-Bench-Real上将Jaccard指数从40.66提高到75.44。此外,ATR能够可靠地引导精确率-召回率工作点,支持具有异构策略要求的部署场景。代码和检查点可在提供的https链接中获取。

英文摘要

The rapid growth of video-based social media has increased users' exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision-recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision-recall operating point, supporting deployment scenarios with heterogeneous policy requirements. Code and checkpoints are provided at https://bruceyg.github.io/ATPO-project-page/ .

发表机构

  • University of Cambridge(剑桥大学)
  • Xiaohongshu Inc.(小红书公司)
  • AntGroup(蚂蚁集团)
  • Tencent Company, China(腾讯公司)
  • University of Bath(巴斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑