arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

K空间特征:用于医学深度伪造检测的频域表示学习

The K-Space Signature: Frequency-Domain Representation Learning for Medical Deepfake Detection

Riccardo Raciti, Francesco Guarnera, Francesco Rundo, Luca Guarnera, Sebastiano Battiato

arXiv 2607.29541首次发表:更新:

AI 中文总结

针对医学深度伪造威胁,提出KSS取证框架,结合3D MLP-Mixer与ArcFace头,在多中心3D MRI数据集上检测准确率超0.99,零样本泛化至未见过的扫描仪数据集,准确率达0.93。

AI 中文摘要

在医学成像领域,生成模型正越来越多地被用于合成逼真数据和扩充有限的数据集。尽管这对隐私保护型数据共享有益,但这些合成图像可能被用于恶意目的,通过生成医学深度伪造产品威胁公共卫生。为应对这一威胁,我们提出了K空间特征(KSS),这是一种在频域内分离硬件与生成痕迹的新型取证框架。通过将分析转移至频域,KSS在对数功率谱密度(Log-PSD)空间中减去经验全局解剖先验,从而抑制宏观解剖变异。为有效处理这些全局分布的频谱伪影,同时避免卷积神经网络固有的局部空间偏差,我们将KSS表示与配备ArcFace度量学习头的新型3D MLP-Mixer架构相结合。在多中心3D MRI数据集上开展的大量实验表明,该组合方法实现了出色的检测性能,在多生成器合成数据集上的准确率和ROC-AUC均超过0.99。此外,该框架展现出强大的零样本泛化能力,在完全未见过的扫描仪采集的独立数据集上仍保持强判别力,准确率最高达0.93。为确保完全可复现,完整源代码和预训练模型将在论文接收后公开。

英文摘要

In medical imaging, generative models are increasingly deployed to synthesize realistic data and augment limited datasets. Unfortunately, while beneficial for privacy-preserving data sharing, these synthesized images can be repurposed for malicious intents, threatening public health through the creation of Medical Deepfakes. To address this threat, we introduce the K-Space Signature (KSS), a novel forensic framework that isolates hardware and generative traces within the spectral domain. By shifting analysis to the frequency domain, the KSS suppresses macroscopic anatomical variance by subtracting an empirical global anatomical prior computed in the Logarithmic Power Spectral Density (Log-PSD) space. To effectively process these globally distributed spectral artifacts without the local spatial bias inherent to Convolutional Neural Networks, we pair the KSS representation with a novel 3D MLP-Mixer architecture equipped with an ArcFace metric-learning head. Extensive experiments on multi-center 3D MRI datasets demonstrate that this combined approach achieves exceptional detection performance, exceeding 0.99 Accuracy and ROC-AUC on multi-generator synthetic datasets. Furthermore, the framework exhibits robust zero-shot generalization, maintaining strong discriminative power (up to 0.93 Accuracy) on independent datasets acquired from entirely unseen scanners. To ensure full reproducibility, the complete source code and pre-trained models will be made publicly available upon acceptance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑