发表机构
Mila - Quebec Artificial Intelligence Institute; Polytechnique de Montréal; Research Center for Information Technology Innovation, Academia Sinica; McGill University(米拉-魁北克人工智能研究所; 蒙特利尔理工学院; 中央研究院资讯科技创新研究中心; 麦吉尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对Transformer模型权重最低有效位可被用作隐蔽信道的问题,提出GrayShield训练后零数据净化方法,用格雷码覆盖该信道,实现近随机比特精度,且精度影响低于1%。
AI 中文摘要
BERT和Vision Transformer(ViT)等Transformer模型通过密集参数化的注意力骨干网络实现了强大的性能。然而,其32位浮点权重的最低有效位(LSBs)可能被滥用为隐蔽信道,以隐藏恶意载荷,对AI模型供应链构成严重威胁。我们提出GrayShield(简称GS),一种轻量级、训练后、零数据的净化方法,该方法用格雷码引导的低跳变序列完全替换声明的尾数LSB信道。完整的与载荷无关的覆盖(无论是密钥化还是公开)使净化后的目标位与嵌入的载荷无关,并使该声明信道具有零容量。格雷码提供覆盖结构,而密钥化的逐张量相位提供模式多样性。在四种Transformer模型预设和两种真实世界恶意载荷上,与七种训练后防御方法进行基准测试,GrayShield在五种实现的攻击变体下保持了低于1%的精度影响,并实现了49.96±0.66个百分点的恢复降低(RR)。由于防御前的恢复率实际上为100%,接近50个百分点的RR对应于净化后比特精度处于随机水平。其主要经验优势是稳定的接近随机的净化效果,且权重分布偏移显著小于评估的接近随机基线PatternMask(PM)和训练后量化(PTQ)。
英文摘要
Transformer models such as BERT and Vision Transformer~(ViT) achieve strong performance via densely parameterized attention backbones. However, the least significant bits~(LSBs) of their 32-bit floating-point weights can be abused as covert channels to conceal malicious payloads, posing a serious threat to the AI model supply chain. We propose \GS (\GSabbr), a lightweight, post-training, zero-data sanitization method that completely replaces the declared mantissa-LSB channel with a Gray-code-guided low-transition sequence. Complete payload-independent overwrite, whether keyed or public, makes the sanitized target bits independent of the embedded payload and gives that declared channel zero capacity. Gray coding supplies overwrite structure, while a keyed per-tensor phase supplies pattern diversity. Benchmarked against seven post-training defenses on four Transformer model presets and two real-world malware payloads, \GSabbr maintains sub-$1\%$ accuracy impact and achieves $49.96\pm0.66$ percentage-point Recovery Reduction (RR) under five implemented attacker variants. Because pre-defense recovery is effectively $100\%$, RR near 50 percentage points corresponds to post-sanitization bit accuracy at binary chance. Its main empirical advantage is stable near-chance sanitization with substantially smaller weight-distribution shift than the evaluated near-chance baselines PatternMask (PM) and Post-Training Quantization (PTQ).
Comments17 pages, 6 figures, conference