发表机构
Universidade da Beira Interior; Shri Guru Buddhi Swami College; Swami Ramanand Tirtha Marathwada University(贝拉内斯大学; 什里·古鲁·布迪·斯瓦米学院; 斯瓦米拉曼南德·蒂尔塔马哈拉施特拉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对深度伪造图像检测的三种范式缺乏通用评估的问题,提出VendorBench-100基准测试,用统一框架评估36个模型,强调现实场景,通过MCC排名,发现ROC-AUC与MCC有差异,还发布评估框架和结果助力后续研究。
AI 中文摘要
深度伪造图像检测目前由三种根本不同的范式提供服务:商业应用程序编程接口(APIs)、零样本视觉语言模型(LLMs)和开源检测器。尽管它们被广泛使用,但这些范式很少在通用协议下进行评估,难以直接比较。我们引入了VendorBench-100,这是一个跨范式基准测试,使用单个对抗性的100图像语料库、统一的输出模式和通用评估框架来评估36个代表性模型。为确保在语料库有意的类别不平衡下进行可靠评估,模型主要根据马修斯相关系数(MCC)排名,同时报告ROC-AUC作为排名能力的阈值无关度量。VendorBench-100通过精心策划的八个边缘案例家族分类法强调具有挑战性的现实世界场景。我们的评估表明,商业APIs实现了最强的中位数性能,其次是视觉LLMs和开源检测器。更重要的是,我们发现排名能力(ROC-AUC)和操作点质量(MCC)之间存在一致的差异。我们发布了完整的评估框架和基准测试结果以支持可重复的未来研究。
英文摘要
Deepfake image detection is served by three fundamentally different paradigms - commercial APIs, zero-shot vision-language models (LLMs), and open-source detectors - that are rarely evaluated under a common protocol, making direct comparison difficult. We introduce VendorBench-100, a cross-paradigm benchmark that evaluates 36 representative models using a single adversarial 100-image corpus, a unified output schema, and a common evaluation framework. Models are ranked primarily by the Matthews correlation coefficient (MCC), with ROC-AUC reported as a threshold-independent measure of ranking ability. Rather than maximizing size, it emphasizes real-world difficulty through a taxonomy of eight edge-case families such as face swaps, text-to-video stills, AI photo edits, avatar compositing, opaque-provenance images, and compressed research frames. Commercial APIs achieve the strongest median performance, followed by vision LLMs and open-source detectors, though individual open-source models remain competitive with the best LLMs. Across all 36 models, MCC and ROC-AUC are strongly correlated (Pearson r ~ 0.86); the more consequential finding is narrower and one-directional: a subset of otherwise strong rankers are miscalibrated at their shipped default threshold, so a high ROC-AUC can overstate real-world deployability. Separately, raw accuracy and F1 are unreliable on this corpus's imbalanced class split, since a model that predicts "fake" indiscriminately scores deceptively well on both while offering no real discriminative skill. No single metric is safe in isolation: MCC and specificity should always accompany ROC-AUC and accuracy. We release the complete evaluation framework and results. Code and data: https://github.com/sharayu-20/vendorbench-100
CommentsFixed scoring-normalization error in Neural Defend/TruthScan results, revising the central finding: MCC and ROC-AUC are strongly correlated (r ~ 0.86) overall, with only a specific subset showing one-directional threshold miscalibration. Abstract, tables, Discussion, Conclusion updated