arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IMFD:基于指令的大型视觉语言模型的端到端多脸伪造检测

IMFD: End-to-end Multi-Face Forgery Detection through Instruction-based Large Vision-Language Models

Dasom Choi, Sangjun Moon, Hyeongchan Im, Jaeeon Park, Jingun Kwon, Hidetaka Kamigaito, Taro Watanabe, Manabu Okumura

arXiv 2609.19693首次发表:更新:

发表机构

Chungnam National University; Institute of Science Tokyo; Nara Institute of Science and Technology (NAIST)(忠南大学; 东京科学大学; 奈良先端科学技术大学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对传统多脸伪造检测忽略上下文与人脸关系的问题,提出基于指令的大型视觉语言模型IMFD,端到端联合定位与检测,将人脸框融入指令,显著提升性能。

AI 中文摘要

深度伪造的迅速增加因其在社交媒体上的传播而引发了重大关注。传统的多脸伪造检测器会裁剪并独立验证每张人脸,忽略了背景上下文和人脸间的关系,这往往导致次优的性能。为克服这些限制,我们利用了基于指令的大型视觉语言模型(LVLMs),这些模型能够解释整个图像并遵循复杂的文本指令。我们提出了一种简单而有效的单阶段多脸伪造检测器,称为IMFD(基于指令的多脸伪造检测器),它以端到端方式训练,联合定位人脸并预测每张人脸的伪造标签。IMFD并未将人脸框预测仅视为联合目标,而是将预测的人脸边界框明确整合到指令中,作为增强指令接地和伪造检测的视觉线索。为支持IMFD的训练和评估,我们将现有的多脸伪造数据集转换为基于指令的格式。实验结果和分析表明,IMFD通过将人脸边界框整合到指令中,改善了多脸伪造检测,并持续优于各种最先进的方法。

英文摘要

The rapid increase of deepfakes has raised significant concerns due to their spread on social media. Traditional multi-face forgery detectors crop and verify each face independently, ignoring background context and inter-face relationships, which often yields suboptimal performance. To overcome these limitations, we leverage instruction-based Large Vision-Language Models (LVLMs), which can interpret entire images and follow complex textual instructions. We propose a simple yet effective single-stage multi-face forgery detector, called IMFD (Instruction-based Multi-face Forgery Detector), which is trained end-to-end to jointly localize faces and predict per-face forgery labels. Rather than treating face box prediction only as a joint objective, IMFD explicitly integrates predicted face bounding boxes into the instruction as visual cues that enhance instruction grounding and forgery detection. To support the training and evaluation of IMFD, we convert existing multi-face forgery datasets into an instruction-based format. Experimental results and analyses show that IMFD improves multi-face forgery detection by integrating face bounding boxes into the instruction, and consistently outperforms various state-of-the-art methods.

Comments8 pages, 5 figures, 5 tables. Accepted to Findings of AACL-IJCNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑