AI 中文总结
研究针对仇恨性视觉错觉检测难题,提出自适应视图检索方法,将其公式化为感知检索问题,通过检索并校准框架组装视图库,该方法在多方面超越基线和其他方法,表明多模态审核需先恢复隐藏含义再判断是否有害。
AI 中文摘要
仇恨性视觉错觉暴露了当前多模态安全系统中的严重差距。对于原始视图仇恨性错觉,先前工作表明六个审核分类器的准确率最高为20.9%至24.5%,九个先进的视觉语言模型在有错觉感知提示的情况下仍处于或低于10.2%,导致大多数隐藏仇恨未被检测到。我们将隐藏仇恨性错觉检测公式化为感知检索问题并提出自适应视图检索。这个检索并校准的框架为图像和隐藏消息模板组装一个互补视图库,自适应选择信任的视图,检索隐藏消息身份并校准恢复的证据是否有害。在使用冻结的CLIP编码器的HatefulIllusion上,自适应视图检索在保留的测试分割上达到93.2%的平衡准确率。它在仇恨俚语、仇恨符号和可见性水平方面大大优于原始视图基线和固定单变换滤波器。相同设计还超越了官方微调的CLIP基线,在IllusionMNIST、IllusionFashionMNIST和IllusionAnimals上匹配或超过人类性能,并在SemVink协议下的HC-Bench上优于缩小预处理。这些结果共同表明,强大的多模态审核需要在判断隐藏含义是否有害之前恢复其隐藏意义。
英文摘要
Hateful optical illusions expose a serious gap in current multimodal safety systems. On original-view hateful illusions, previous work shows that six moderation classifiers achieve at most 20.9 to 24.5% accuracy and nine state-of-the-art VLMs remain at or below 10.2% with illusion-aware prompting, leaving most hidden hate undetected. We formulate hidden hateful illusion detection as a perceptual retrieval problem and propose Adaptive View Retrieval. This retrieve-and-calibrate framework assembles a complementary view bank for the image and hidden-message templates, adaptively selects which views to trust, retrieves hidden-message identities, and calibrates whether the recovered evidence is harmful. On HatefulIllusion with a frozen CLIP encoder, Adaptive View Retrieval reaches 93.2% balanced accuracy on the held-out test split. It substantially outperforms original-view baselines and fixed single-transform filters across hate slangs, hate symbols, and visibility levels. The same design also surpasses official fine-tuned CLIP baselines, matches or exceeds human performance on IllusionMNIST, IllusionFashionMNIST, and IllusionAnimals, and outperforms zoom-out preprocessing on HC-Bench under the SemVink protocol. Together, these results show that robust multimodal moderation requires recovering hidden meaning before deciding whether it is harmful.