arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31661cs.CV

ForensicZoom:基于多模态大语言模型的自适应视觉检查用于工业级人脸伪造检测

ForensicZoom: Adaptive Visual Inspection with Multimodal LLMs for Industrial-Grade Face Forgery Detection

Hang Zhou, Yiming Tang, Kun Yu, Qian Zhu, Minghao Li, Weigao Wen

首次发表
浏览论文内容

中文总结 AI 辅助

针对人脸伪造检测中细微伪影难以捕捉及计算效率问题,提出ForensicZoom框架,通过NEED_ZOOM机制自适应放大可疑区域,结合奖励塑形与归因优化,在工业数据上实现97%以上TPR@0.1%FPR,兼顾准确性与可解释性。

中文摘要 AI 辅助

可靠的人脸伪造检测对于在线身份验证系统的安全至关重要,因为漏检会危及安全,而过多的误报会干扰合法用户。专门的取证检测器实现了强大的检测性能,但可解释性有限,而多模态大语言模型(MLLMs)提供了强大的语义理解和可解释的推理能力,但在人脸伪造检测方面仍然明显较弱。我们认为一个关键的局限性在于视觉证据的获取方式:细微的取证伪影在标准分辨率下可能无法得到良好表示,而将所有案例统一以更高分辨率处理在计算上效率低下。因此,我们引入了ForensicZoom,一个用于自适应视觉检查的工业级MLLM框架。ForensicZoom首先为通用MLLM配备取证感知的视觉表示,并将语言模型与这些特征对齐。其核心机制NEED_ZOOM使模型能够在初始证据不足时自主请求可疑区域的放大视图,将固定次数的分类转变为自适应多轮取证推理。缩放行为通过奖励塑形学习,该机制在检测准确性与不必要的视觉检查之间取得平衡,将额外计算集中在困难案例上。最后的归因优化阶段在保持检测性能的同时改进自然语言取证报告。在大型工业身份验证数据上,ForensicZoom在0.1%假阳性率(FPR)下实现了超过97%的真阳性率(TPR),大幅优于专门的检测器和现有的基于MLLM的方法,同时生成可操作的取证归因。这些结果表明,ForensicZoom可以为基于MLLM的人脸伪造检测提供一条通往准确、可解释且可扩展的有效路径。

英文摘要

Reliable face forgery detection is critical to the security of online identity verification systems, where missed attacks compromise security and excessive false positives disrupt legitimate users. Specialized forensic detectors achieve strong detection performance but provide limited interpretability, while multimodal large language models (MLLMs) offer strong semantic understanding and interpretable reasoning yet remain substantially weaker for face forgery detection. We argue that a key limitation lies in how visual evidence is acquired: subtle forensic artifacts may be poorly represented at standard resolution, while uniformly processing all cases at higher resolution is computationally inefficient. We therefore introduce ForensicZoom, an industrial-grade MLLM framework for adaptive visual inspection. ForensicZoom first equips a general-purpose MLLM with forensic-aware visual representations and aligns the language model with these features. Its central mechanism, NEED_ZOOM, enables the model to autonomously request magnified views of suspicious regions when the initial evidence is insufficient, turning fixed-pass classification into adaptive multi-round forensic reasoning. The zoom behavior is learned through reward shaping that balances detection accuracy with unnecessary visual inspection, concentrating additional computation on difficult cases. A final attribution optimization stage improves natural-language forensic reports while preserving detection performance. On large-scale industrial identity verification data, ForensicZoom achieves over 97% TPR at 0.1% FPR, substantially outperforming both specialized detectors and existing MLLM-based methods while producing actionable forensic attributions. These results demonstrate that ForensicZoom can provide an effective path toward accurate, interpretable, and scalable MLLM-based face forgery detection.

↑