语言模型中量化触发的后门:跨量化器可迁移性与验证-部署差距
Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap
浏览论文内容
中文总结 AI 辅助
该研究发现仅源精度审计无法排除语言模型的量化触发后门,提出量化行为等价类理论,在多语言模型中实现量化后激活的后门攻击,证明攻击持久性与量化方案及模型架构相关。
中文摘要 AI 辅助
训练后量化通常被视为大型语言模型边缘部署的语义中性优化。当对全精度源检查点进行评估后,在下游应用量化却未进行同等的重新评估时,该工作流会产生结构性的验证-部署差距:由于量化是参数空间上的多对一映射,源精度认证无法保证部署配置中的行为等价性。我们通过量化行为等价类(QBECs)将此差距形式化,并证明QBEC成员关系不意味着行为等价性,为量化触发的后门攻击提供理论基础。基于三阶段对抗微调框架,我们将潜在恶意载荷嵌入到通过评估所用源精度检查的模型中,这些模型在INT8或4位压缩时会激活目标对抗行为。我们在两个具有操作动机的场景(战术机器翻译和政治内容分析)中评估此威胁,将先前工作从仅解码器的因果语言模型扩展到多语言编码器-解码器序列到序列模型。结果显示,被植入后门的翻译模型在修复后的FP16下测得的友敌腐败为零,量化后却出现高达85.02%的反转;配对立场分类器在压缩后测得的意识形态偏移高达ΔBias=0.33。跨量化器可迁移性分析进一步表明,攻击的持久性因量化方案和模型架构而异,而非仅由名义位宽决定。这些发现表明,仅源精度审计无法排除量化触发的行为,可信边缘AI的行为认证必须包含最终部署配置。
英文摘要
Post-training quantization is often treated as a semantically neutral optimization for edge deployment of Large Language Models. When a full-precision source checkpoint is evaluated and quantization is applied downstream without equivalent re-evaluation, this workflow creates a structural validation--deployment gap: because quantization is a many-to-one mapping over parameter space, source-precision certification does not guarantee behavioral equivalence in the deployed configuration. We formalize this gap through Quantization Behavioral Equivalence Classes (QBECs) and prove that QBEC membership does not imply behavioral equivalence, providing a theoretical basis for quantization-triggered backdoor attacks. Building on a three-stage adversarial fine-tuning framework, we embed latent malicious payloads into models that satisfy the source-precision checks used in our evaluation, yet activate targeted adversarial behavior upon INT8 or 4-bit compression. We evaluate this threat in two operationally motivated scenarios, tactical machine translation and political content analysis, extending prior work from decoder-only causal LMs to multilingual encoder-decoder sequence-to-sequence models. Results show that backdoored translation models move from zero measured friend--foe corruption at repaired FP16 to up to 85.02% inversion after quantization, and that a paired stance classifier measures an ideological shift of up to $Δ\mathrm{Bias}=0.33$ upon compression. A cross-quantizer transferability analysis further shows that attack persistence varies across quantization schemes and model architectures, rather than being determined by nominal bit-width alone. These findings demonstrate that source-precision auditing alone does not rule out quantization-triggered behavior and that the final deployed configuration must be included in behavioral certification for trustworthy edge AI.
发表机构
- University of Bologna(博洛尼亚大学)
- Luiss Guido Carli University(路易斯 Guido Carli 大学)
- Live Tech
- University of Salerno(萨勒诺大学)
机构由 AI 辅助整理,请以论文原文为准。