摘要、评判、优化:多模态内容审核的解耦内容理解与策略学习
Summarize, Judge, Refine: Decoupled Content Understanding and Policy Learning for Multimodal Content Moderation
浏览论文内容
中文总结 AI 辅助
提出SJR双模型架构,通过内容模型生成摘要、策略模型分类,并利用GRPO协同训练与文本增强,在误导广告检测中提升F1达23.6%,且可在无真实违规数据下启动新策略。
中文摘要 AI 辅助
传统的内容审核系统将多模态理解与特定策略的分类纠缠在一起,导致每次策略变更都需要重新训练整个流程,并且由于多媒体无法进行有意义的增强,系统还面临标签稀缺的问题。我们提出了摘要-评判-优化(SJR)架构,这是一种通过自然语言接口解耦上述问题的双模型架构:一个多模态内容模型生成结构化的文本摘要,一个纯文本策略模型根据策略定义对这些摘要进行分类。一个迭代的协同训练循环通过GRPO优化内容模型,以生成与策略相关的摘要,而文本空间增强则生成对抗性的摘要变体——这是一种在原始多媒体上不可能实现的增强途径——从而实现少样本策略启动。每个决策都基于人类可读的摘要,作为结构性副产品提供了可解释性。在误导性广告检测中,SJR相对于零样本思维链基线实现了+23.6%的相对非误导F1提升,优于端到端SFT、STaR/RFT和RLFT方法。值得注意的是,一个在零真实违规示例上训练(所有正类数据均为合成生成)的变体,在违规F1上与全数据模型的相对差距在0.2%以内,这表明新策略可以在没有任何真实违规数据的情况下启动。
英文摘要
Content moderation systems traditionally entangle multimodal understanding with policy-specific classification, requiring full pipeline retraining for every policy change and suffering from label scarcity since multimedia cannot be meaningfully augmented. We propose Summarize-Judge-Refine (SJR), a two-model architecture that decouples these concerns via a natural language interface: a multimodal Content Model produces structured text summaries, and a text-only Policy Model classifies them against policy definitions. An iterative co-training loop refines the Content Model via GRPO to produce policy-relevant summaries, while text-space augmentation generates adversarial summary variants---an augmentation pathway impossible on raw multimedia---enabling few-shot policy bootstrap. Every decision is grounded in a human-readable summary, providing interpretability as a structural byproduct. On misleading advertisement detection, SJR achieves +23.6\% relative non-misleading F1 over a zero-shot chain-of-thought baseline, outperforming end-to-end SFT, STaR/RFT, and RLFT. Notably, a variant trained on zero real violating examples---with all positive-class data synthetically generated---matches the full-data model within 0.2\% relative on violating F1, demonstrating that new policies can launch without any real violation data.
发表机构
- Meta AI
机构由 AI 辅助整理,请以论文原文为准。