发表机构
Tarbiat Modares University; University of Ottawa; Technische Universität Braunschweig(塔比阿特莫达雷斯大学; 渥太华大学; 不伦瑞克工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对语义通信易受对抗扰动攻击的问题,提出TwinViT-DeepJSCC收发机,采用双分支ViT编码、敏感性感知掩蔽及DDIM净化等机制,在CIFAR-100上显著提升PSNR与Top-1准确率,实现鲁棒图像传输。
AI 中文摘要
基于学习的语义通信容易受到在语义编码前或无线信道上引入的对抗性扰动的影响。本文提出了TwinViT-DeepJSCC,一种在固定信道使用预算下运行的预防-纠正型语义图像收发机。两个基于视觉Transformer(ViT)的深度联合信源信道编码(DeepJSCC)分支学习互补的潜在表示,并通过敏感性感知掩蔽加以保护。在接收端,置信度感知融合、盲损坏严重性估计以及信噪比(SNR)-严重性条件化的去噪扩散隐式模型(DDIM)净化可减轻残余损坏,而无需攻击元数据。在加拿大高等研究院100类(CIFAR-100)数据集上的实验考虑了快速梯度符号法(FGSM)、投影梯度下降(PGD)、自然进化策略(NES)和Carlini-Wagner(CW)源域攻击,以及加性高斯白噪声(AWGN)和块平坦瑞利衰落信道上的随机干扰和信道感知对抗波形。在匹配的信道使用和攻击预算下,TwinViT-DeepJSCC在20步PGD下实现了约9.5 dB的最大峰值信噪比(PSNR)增益,在块平坦瑞利衰落下的信道感知波形攻击下实现了约10.8 dB的增益。在PGD下,它还将Top-1准确率比无防御基线提高了约38个百分点,比最强竞争防御提高了13个百分点。消融结果证实了所提出的发射端和接收端机制的互补贡献。
英文摘要
Learning-based semantic communication is vulnerable to adversarial perturbations introduced before semantic encoding or over wireless channels. This paper proposes TwinViT-DeepJSCC, a preventive-corrective semantic image transceiver operating under a fixed channel-use budget. Two Vision Transformer (ViT)-based deep joint source-channel coding (DeepJSCC) branches learn complementary latent representations protected by sensitivity-aware masking. At the receiver, confidence-aware fusion, blind corruption-severity estimation, and signal-to-noise ratio (SNR)-severity-conditioned denoising diffusion implicit model (DDIM) purification mitigate residual corruption without requiring attack metadata. Experiments on the Canadian Institute for Advanced Research 100-class (CIFAR-100) dataset consider fast gradient sign method (FGSM), projected gradient descent (PGD), natural evolution strategies (NES), and Carlini-Wagner (CW) source-domain attacks, as well as random jamming and channel-aware adversarial waveforms over additive white Gaussian noise (AWGN) and block-flat Rayleigh fading. Under matched channel-use and attack budgets, TwinViT-DeepJSCC achieves maximum peak signal-to-noise ratio (PSNR) gains of approximately 9.5 dB under 20-step PGD and 10.8 dB under channel-aware waveform attacks over block-flat Rayleigh fading. Under PGD, it also improves Top-1 accuracy by up to approximately 38 percentage points over the undefended baseline and 13 percentage points over the strongest competing defense. Ablation results confirm the complementary contributions of the proposed transmitter- and receiver-side mechanisms.
Comments13 pages, 7 figures