发表机构
ADA University; MegaSec AI Company(ADA大学; MegaSec人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过受控实验发现,抗压缩深度伪造检测的关键是数据多样性而非频率不变性,普通EfficientNet-B0优于CAFRL,对抗分支无增益,需匹配训练配方并利用编解码器多样性提升鲁棒性。
AI 中文摘要
人们普遍认为频率特征和压缩不变表示学习是抗视频压缩的深度伪造检测的关键。我们通过CAFRL-block-DCT和FFT相位流、压缩级别条件带注意力机制以及对抗(梯度反转)压缩不变性来验证这一点,并报告了一项受控的负面结果。在容量和增强匹配的对照、预注册协议下,在FaceForensics++测试集的每个压缩级别上,针对指定的CAFRL,在多质量数据上的普通EfficientNet-B0均表现更优,在CRF 40时AUC高出3.66个点(配对单种子)。对我们自身负面结果的自审计发现了四个不利于频率假设的缺陷,而修复所有缺陷的预指定重测显示,该缺陷是配方人为因素导致,而非架构故障:基线配方比匹配的 shipped-recipe 变体恢复了3.96个点。频率路径未产生可检测的差异:单独使用时具有判别性(训练后期独立验证AUC为0.91-0.98),但在该融合设置下无边际价值,在4.0 M主干的两种特征宽度下,数据集内压缩对比的所有种子池区间均包含零;在测试的单个保留操作上,公平变体低于普通主干。按指定的对抗分支未增加任何效果,且会降低其自身的条件估计器;在公平配方下该分支未被测试。单遍H.264重编码下的鲁棒性反而来自数据多样性:真实恒定速率因子变体比合成JPEG增强高出7.3个点(单次运行,非重叠区间)。证据来自FaceForensics++系列、GAN时代及单编解码器设置。需在训练配方和容量上进行匹配控制,并在架构选择前通过编解码器多样性获取压缩鲁棒性。
英文摘要
Frequency features and compression-invariant representation learning are widely assumed to be key to deepfake detection that survives video compression. We test this with CAFRL - block-DCT and FFT-phase streams, compression-level-conditioned band attention, and adversarial (gradient-reversal) compression invariance - and report a controlled negative. Under a pre-registered protocol with capacity- and augmentation-matched controls, a plain EfficientNet-B0 on multi-quality data beat CAFRL as specified at every compression level on the FaceForensics++ test split, by 3.66 AUC points at CRF 40 (paired, single seed). A self-audit of our own negative found four defects biased against the frequency hypothesis, and pre-specified re-tests repairing all four showed the deficit to be a recipe artifact, not an architecture failure: the baseline recipe recovered 3.96 points over the matching shipped-recipe variant. The frequency path made no detectable difference: discriminative alone (standalone validation AUC 0.91-0.98 late in training) but of no marginal value under this fusion, at two feature widths of one 4.0 M trunk, every seed-pooled interval for the intra-dataset compression contrasts including zero; on the single held-out manipulation tested, the fair variants sat below the plain backbone. The adversarial branch, as specified, added nothing and degraded its own conditioning estimator; at the fair recipe it is untested. Robustness under single-pass H.264 re-encoding came instead from data diversity: real constant-rate-factor variants beat synthetic JPEG augmentation by 7.3 points (single runs, non-overlapping intervals). The evidence is FaceForensics++-family, GAN-era and single-codec. Match controls on training recipe as well as capacity, and buy compression robustness with codec diversity before architecture.
Comments38 pages, 4 figures. Submitted to IEEE Access. Code, per-video predictions and a claims-to-artifacts map: https://github.com/Capta1n-n9m0/cafrl-negative (v1.0.0); archival deposit: https://doi.org/10.5281/zenodo.22032529