基于文本引导扩散模型的胸部X射线图像对抗攻击
Text-Guided Diffusion-Based Adversarial Attacks on Chest X-Ray Images
浏览论文内容
中文总结 AI 辅助
该研究提出文本引导扩散对抗框架攻击胸部X射线分类模型,在多架构上验证其能显著降低分类性能,同时发现人机判读存在重要差异,凸显需扩展医疗AI鲁棒性评估方式。
中文摘要 AI 辅助
随着人工智能越来越多地被应用于胸部X射线(CXR)的判读、分诊及临床决策支持,了解其对抗操纵的脆弱性对于安全部署至关重要。然而,现有的鲁棒性评估主要依赖于像素空间攻击,这类攻击引入的扰动在数值上受到限制,但可能无法代表合理的放射学变异。这一局限在多疾病CXR分类中尤为重要,在此场景下,模型需同时评估多种重叠的病理状况,而对抗性失败可能会改变多项诊断预测。我们提出了一种文本引导的扩散对抗框架,该框架在保持扩散生成器和目标分类器冻结的同时,优化可学习的文本条件,从而通过学习到的图像先验而非直接的像素操作实现对抗性生成。我们在二元肺不张分类和多疾病CXR分类的多种分类器架构上对该框架进行了评估,并将其与FGSM、PGD和Carlini-Wagner攻击进行了比较。我们的方法始终导致分类器性能出现最大幅度的下降,在二元分类中将AUROC降至0.3885-0.5646,在多疾病场景下降至0.4441-0.4878,同时实现了更优的图像保真度(SSIM为0.9080,LPIPS为0.1670,FID为51.23)。重要的是,尽管模型预测发生了显著变化,但95.9%的二元对抗图像和73.8%的多疾病对抗图像的临床医生判读仍保持不变。这些发现揭示了人类与机器判读之间存在临床意义上的重要差异,并表明需要将医疗AI的鲁棒性评估从传统的像素空间攻击扩展到能够在视觉和临床合理的图像变异下暴露失败的生成式威胁模型。
英文摘要
As artificial intelligence is increasingly integrated into chest X-ray (CXR) interpretation, triage, and clinical decision support, understanding its vulnerability to adversarial manipulation is critical for safe deployment. Existing robustness evaluations, however, predominantly rely on pixel-space attacks that introduce numerically constrained perturbations but may not represent plausible radiographic variation. This limitation is particularly important in multi-disease CXR classification, where models simultaneously evaluate multiple overlapping pathologies and adversarial failures may alter several diagnostic predictions. We propose a text-guided diffusion-based adversarial framework that optimizes learnable text conditioning while keeping the diffusion generator and target classifier frozen, enabling adversarial generation through a learned image prior rather than direct pixel manipulation. We evaluate the framework across multiple classifier architectures in both binary atelectasis and multi-disease CXR classification and compare it with FGSM, PGD, and Carlini-Wagner attacks. Our approach consistently produced the greatest degradation in classifier performance, reducing AUROC to 0.3885-0.5646 in binary classification and 0.4441-0.4878 in the multi-disease setting, while achieving superior image fidelity (SSIM 0.9080, LPIPS 0.1670, FID 51.23). Importantly, clinician interpretation remained unchanged for 95.9% of binary and 73.8% of multi-disease adversarial images despite substantial changes in model predictions. These findings reveal a clinically important discrepancy between human and machine interpretation and demonstrate the need to extend medical AI robustness evaluation beyond conventional pixel-space attacks toward generative threat models that can expose failures under visually and clinically plausible image variations.
发表机构
- The Johns Hopkins University(约翰斯·霍普金斯大学)
- Manipal Institute of Technology(马尼帕尔理工学院)
- Manipal Academy of Higher Education(马尼帕尔高等教育学院)
- The Johns Hopkins Hospital(约翰斯·霍普金斯医院)
- Columbia University Irving Medical Center(哥伦比亚大学欧文医学中心)
机构由 AI 辅助整理,请以论文原文为准。