用于人脸换脸检测的相机噪声残差:冗余而非互补,及原因
Camera-Noise Residuals for Face-Swap Detection: Redundant, Not Complementary, and Why
浏览论文内容
中文总结 AI 辅助
该研究在FaceForensics++上发现,相机噪声残差与RGB外观骨干网络融合用于人脸换脸检测时,噪声残差为冗余而非互补,因InstanceNorm消除了其统计信号,且融合未优于单独RGB。
中文摘要 AI 辅助
将学习得到的相机噪声指纹与RGB外观骨干网络融合,是一种颇具吸引力的生成器无关深度伪造检测路线,因为噪声残差基于图像形成物理原理,而非特定生成器的纹理统计特征。我们在FaceForensics++数据集上,测试Noiseprint++残差通道是否为RGB Xception骨干网络提供人脸换脸检测的互补信息。三模型消融实验(仅RGB、仅残差、后期融合)显示,融合效果未优于单独使用RGB,且单独使用残差分支的表现接近随机水平。七层级瓶颈诊断定位了原因:噪声图确实携带判别信号,但属于统计信号,由残差的样本一阶矩(均值)、二阶矩(方差、能量)承载;而遵循TruFor模板置于噪声分支输入的样本级InstanceNorm层,恰好将这些矩标准化消除(五折交叉验证AUC从0.747降至0.554)。上下文裁剪对照排除了裁剪几何的影响,两个移除瓶颈的固定融合变体虽恢复了统计信号,但在所有数据集上仍未能击败RGB。我们得出结论:在该操作分布下,噪声残差与RGB是冗余而非互补的,并为采用噪声残差融合进行人脸换脸检测的从业者提供了具体指导。
英文摘要
Fusing a learned camera-noise fingerprint with an RGB appearance backbone is an appealing route to generator-independent deepfake detection, because the noise residual is grounded in image-formation physics rather than in the texture statistics of a particular generator. We test, on FaceForensics++, whether a Noiseprint++ residual channel carries information \emph{complementary} to an RGB Xception backbone for face-swap detection. A three-model ablation (RGB-only, residual-only, late-fusion) shows that fusion does not improve over RGB alone and that the residual branch alone is near chance. A seven-level bottleneck diagnostic localizes the cause: the noise maps do carry a discriminative signal, but it is statistical---carried by the per-sample first and second moments (mean, variance, energy) of the residual---and the per-sample \texttt{InstanceNorm} layer placed at the noise-branch input, following the TruFor template, standardizes exactly those moments away (five-fold cross-validated AUC drops from $0.747$ to $0.554$). A context-crop control rules out cropping geometry, and two fixed-fusion variants that remove the bottleneck recover the statistical signal yet still fail to beat RGB on every dataset. We conclude that, on this manipulation distribution, the noise residual is redundant with RGB rather than complementary, and we give concrete guidance for practitioners adopting noise-residual fusion for face-swap detection.
发表机构
- Innopolis University(因诺波利斯大学)
机构由 AI 辅助整理,请以论文原文为准。