arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22094cs.CLcs.LG

摘要、评判、优化:多模态内容审核的解耦内容理解与策略学习

Summarize, Judge, Refine: Decoupled Content Understanding and Policy Learning for Multimodal Content Moderation

Zeeshan Ahmed, Yang Qin, Hanqing Huang

首次发表
浏览论文内容

中文总结 AI 辅助

提出SJR双模型架构,通过内容模型生成摘要、策略模型分类,并利用GRPO协同训练与文本增强,在误导广告检测中提升F1达23.6%,且可在无真实违规数据下启动新策略。

中文摘要 AI 辅助

传统的内容审核系统将多模态理解与特定策略的分类纠缠在一起,导致每次策略变更都需要重新训练整个流程,并且由于多媒体无法进行有意义的增强,系统还面临标签稀缺的问题。我们提出了摘要-评判-优化(SJR)架构,这是一种通过自然语言接口解耦上述问题的双模型架构:一个多模态内容模型生成结构化的文本摘要,一个纯文本策略模型根据策略定义对这些摘要进行分类。一个迭代的协同训练循环通过GRPO优化内容模型,以生成与策略相关的摘要,而文本空间增强则生成对抗性的摘要变体——这是一种在原始多媒体上不可能实现的增强途径——从而实现少样本策略启动。每个决策都基于人类可读的摘要,作为结构性副产品提供了可解释性。在误导性广告检测中,SJR相对于零样本思维链基线实现了+23.6%的相对非误导F1提升,优于端到端SFT、STaR/RFT和RLFT方法。值得注意的是,一个在零真实违规示例上训练(所有正类数据均为合成生成)的变体,在违规F1上与全数据模型的相对差距在0.2%以内,这表明新策略可以在没有任何真实违规数据的情况下启动。

英文摘要

Content moderation systems traditionally entangle multimodal understanding with policy-specific classification, requiring full pipeline retraining for every policy change and suffering from label scarcity since multimedia cannot be meaningfully augmented. We propose Summarize-Judge-Refine (SJR), a two-model architecture that decouples these concerns via a natural language interface: a multimodal Content Model produces structured text summaries, and a text-only Policy Model classifies them against policy definitions. An iterative co-training loop refines the Content Model via GRPO to produce policy-relevant summaries, while text-space augmentation generates adversarial summary variants---an augmentation pathway impossible on raw multimedia---enabling few-shot policy bootstrap. Every decision is grounded in a human-readable summary, providing interpretability as a structural byproduct. On misleading advertisement detection, SJR achieves +23.6\% relative non-misleading F1 over a zero-shot chain-of-thought baseline, outperforming end-to-end SFT, STaR/RFT, and RLFT. Notably, a variant trained on zero real violating examples---with all positive-class data synthetically generated---matches the full-data model within 0.2\% relative on violating F1, demonstrating that new policies can launch without any real violation data.

发表机构

  • Meta AI

机构由 AI 辅助整理,请以论文原文为准。

↑