arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

什么改善了多模态虚假信息检测?来自大规模实证研究的答案

What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study

Akshit Sharma, Prashant W. Patil

arXiv 2609.30402首次发表:更新:

发表机构

CVPR Lab, MFSDSAI; Indian Institute of Technology Guwahati(CVPR实验室,MFSDSAI; 印度理工学院古瓦哈提分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

通过超3,375次实验的大规模实证研究,系统评估多模态虚假信息检测中的设计选择,提炼出关键指导建议,为构建更可靠检测系统提供依据。

AI 中文摘要

多模态虚假信息日益被精心制作,通过将文本声明与看似“证明”该声明的图像配对,使其看起来更具说服力。然而在实践中,构建有效的检测器往往取决于一小部分设计选择,而这些选择很少以受控方式进行检验。在本文中,我们对多模态虚假信息检测的设计选择进行了大规模研究,进行了超过3,375次实验,涵盖三个基准数据集以及广泛的预训练视觉和语言骨干网络。通过系统性比较和有针对性的鲁棒性分析,我们提炼出实用的指导建议,说明哪些设计选择有帮助、它们何时会无声地失败,以及流程的哪些方面最能影响模型行为,回答了4个关键研究问题(RQs)。我们旨在为设计更强大、更可靠的多模态虚假信息检测系统提供可靠基础,从而为更广泛的研究社区做出贡献。

英文摘要

Multimodal misinformation is increasingly crafted to look convincing by pairing a textual claim with an image that appears to "prove" it. Yet in practice, building effective detectors often hinges on a small set of design choices that are rarely examined in a controlled way. In this paper, we conduct a large-scale study of multimodal design choices for misinformation detection with over 3,375 experiments- spanning three benchmark datasets and a broad range of pre-trained vision and language backbones. Through systematic comparisons and targeted robustness analyses, we distill practical guidance on which design choices help, when do they fail silently, and what aspects of the pipeline most strongly shape model behavior, answering 4 key Research Questions (RQs). We aim to provide a reliable foundation for designing stronger and more dependable multimodal misinformation detection systems, thus contributing to the broader research community.

CommentsAccepted at the Tenth Widening NLP Workshop (WiNLP), co-located with EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑