持续演进的深度伪造检测:动态检测系统的架构与公开基准评估
Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System
浏览论文内容
中文总结 AI 辅助
研究针对深度伪造检测器在现实与学术基准表现差异的问题,提出通过Bittensor SN34训练的BitMind Forensics,在多数据集评估中表现优异,能持续演进提升检测能力,且评估工具公开可独立验证。
中文摘要 AI 辅助
在学术基准测试中表现近乎完美的深度伪造检测器,在面对现实世界内容时效果不佳:近期的实际应用评估报告显示,最先进的开源模型的AUC下降了45%-50%。我们认为这种差距是结构性的,因为静态检测器是针对不断变化的生成前沿进行一次性训练的。我们提出了BitMind Forensics(BMF),它通过Bittensor SN34进行训练,这是一个开放的对抗性竞赛,不断更新训练分布。我们在19个公共数据集上评估了一个包含图像、通用视频和人类视频检查点的过时导出:标准的面部交换套件(FaceForensics++、Celeb-DF v1/v2/++、DFDC、DFD、UADFV、DF40)以及近期的实际应用和人工智能生成媒体基准(Sumsub、Deepfake-Eval-2024、WildRF、Community Forensics、AIGCDetectBench、GenImage、AI-GenBench、AIGIBench、RAID、GenVidBench、GenVideo-100K)。BMF在Sumsub的原始图像上达到了0.936的AUC,在其完整的四条件操作组(140万张图像)上的合并AUC为0.872,在扰动下保持稳健(JPEG为0.855,下采样为0.799),而GPEN增强提高了检测效果(0.996)。在Deepfake-Eval-2024上,它在图像上与最佳商业检测器匹配(0.915对0.90),在视频上超过了它(0.822对0.79),远远高于最佳开源检测器(0.56和0.63)。它在一个21生成器的人工智能图像面板上达到了0.991的AUC,在GenVidBench上达到了0.918,并且在DFDC(0.947对0.843)和Celeb-DF v2(0.9985对0.956)上超过了经过FF++训练的前沿,这两个数据集都经过了污染审核,在Celeb-DF++上具有统计均等性。在一项时间研究中,连续的过时导出在静态基线训练中不存在的生成器的保留媒体上有所改进(图像从0.842提高到0.902;视频从0.864提高到0.936)。我们用于评估的工具是公开的,在发布时,生产API提供了经过精确评估的快照以供独立验证。
英文摘要
Deepfake detectors that achieve near-perfect scores on academic benchmarks collapse on real-world content: recent in-the-wild evaluations report AUC drops of 45-50% for state-of-the-art open-source models. We argue this gap is structural: static detectors are trained once against a moving generative frontier. We present BitMind Forensics (BMF), trained through Bittensor SN34, an open adversarial competition that continually refreshes the training distribution. We evaluate one dated export comprising image, general-video, and human-video checkpoints across nineteen public datasets: the canonical face-swap suites (FaceForensics++, Celeb-DF v1/v2/++, DFDC, DFD, UADFV, DF40) and recent in-the-wild and AI-generated-media benchmarks (Sumsub, Deepfake-Eval-2024, WildRF, Community Forensics, AIGCDetectBench, GenImage, AI-GenBench, AIGIBench, RAID, GenVidBench, GenVideo-100K). BMF reaches 0.936 AUC on Sumsub's original images and 0.872 pooled AUC over its full four-condition manipulation battery (1.4M images), staying robust under perturbation (0.855 JPEG, 0.799 downscaled), while GPEN enhancement improves detection (0.996). On Deepfake-Eval-2024, it matches the best commercial detector on images (0.915 vs 0.90) and exceeds it on video (0.822 vs 0.79), far above the best open-source detectors (0.56 and 0.63). It reaches 0.991 AUC on a 21-generator AI-image panel and 0.918 on GenVidBench, and exceeds the FF++-trained frontier on DFDC (0.947 vs 0.843) and Celeb-DF v2 (0.9985 vs 0.956), both contamination-audited, with statistical parity on Celeb-DF++. In a temporal study, successive dated exports improve on held-out media from generators absent from the static baseline's training (image 0.842 to 0.902; video 0.864 to 0.936). Our evaluation harness is public, and at publication the production API serves the exact evaluated snapshot for independent verification.
发表机构
- BitMind(比特思维)
机构由 AI 辅助整理,请以论文原文为准。