arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于自适应Tversky策略优化的可控多标签视频安全检测

Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization

Guangyu Yang, Jingbiao Mei, Mingsheng Sun, Jinghong Chen, Yingtong Bu, Pengda Qin, Da Chen, Bill Byrne

arXiv 2610.02019首次发表:更新:

发表机构

University of Cambridge; Xiaohongshu Inc.; AntGroup; Tencent Company, China; University of Bath(剑桥大学; 小红书公司; 蚂蚁集团; 腾讯公司; 巴斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有视频安全检测忽视多标签特性和静态训练目标的问题,提出自适应Tversky策略优化(ATPO)框架,通过动态调整假阳性和假阴性惩罚实现可控精确率-召回率权衡,在SafeWatch-Bench上显著提升多标签性能。

AI 中文摘要

视频社交媒体的快速增长增加了用户接触有害内容的风险,因此需要可靠的自动化视频安全检测系统。尽管近期视觉语言模型(VLMs)展现出强大的视频理解能力,但现有的有害视频检测系统存在两个关键局限:它们通常将安全检测简化为二分类问题,忽视了不安全视频固有的多标签特性;同时,它们依赖静态训练目标,无法支持可控的精确率-召回率权衡,而不同审核流程和不安全类别可能对工作点有不同的需求。为解决这些问题,我们提出了自适应Tversky策略优化(ATPO),一种用于多标签视频安全检测(Multi-VSD)的强化学习框架。ATPO引入了自适应Tversky奖励(ATR),在训练过程中动态调整假阳性和假阴性惩罚,以实现可控的精确率-召回率权衡。在SafeWatch-Bench和XD-Violence上的实验表明,ATPO显著提升了多标签性能,在SafeWatch-Bench-Real上将Jaccard指数从40.66提高到75.44。此外,ATR能够可靠地引导精确率-召回率工作点,支持具有异构策略要求的部署场景。代码和检查点可在提供的https链接中获取。

英文摘要

The rapid growth of video-based social media has increased users' exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision-recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision-recall operating point, supporting deployment scenarios with heterogeneous policy requirements. Code and checkpoints are provided at https://bruceyg.github.io/ATPO-project-page/ .

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑