arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13478cs.CR

BadEngram:针对大语言模型中门控记忆组件的后门攻击

BadEngram: Backdoor Attack on Gated Memory Components in LLMs

发表机构Pillar Security · 富士通欧洲研究院
查看机构详情
  • Pillar Security
  • Fujitsu Research of Europe(富士通欧洲研究院)

机构由 AI 辅助整理,请以论文原文为准。

Ariel Fogel, Omer Hofman, Eilon Cohen, Roman Vainshtein

首次发表
浏览论文内容

中文总结 AI 辅助

提出BadEngram训练后攻击,利用大模型门控记忆组件独立修改参数植入后门,在保持主干不变下实现高攻击成功率,揭示该组件为安全关键部分。

中文摘要 AI 辅助

为了在不按比例增加计算量的情况下扩展开放权重模型的容量,近期语言模型引入了门控参数化记忆模块,这些模块检索学习到的值并将其注入中间表示。尽管具有这些效率优势,此类模块却创造了一个独特的攻击面:其参数可以独立于主干网络进行修改,同时直接塑造其计算过程。我们提出了BadEngram,一种训练后攻击方法,利用该攻击面植入持久且由触发器触发的行为,同时保持传统主干网络权重和执行图不变。我们首先在受控的Engram模型中确立了该攻击的可行性,并从因果角度刻画了其机制,在该模型中BadEngram在触发输入上实现了96.6%的攻击成功率(ASR),同时将匹配的无触发输入上的误激活率限制在0.1%,并保持了99.6%的干净准确率。将检索到的记忆值替换为干净对应值或关闭记忆门控,可将ASR降至至多0.32%,从而确认后门是通过门控记忆通路表达的。随后,我们测试了该漏洞是否扩展到生产规模,即在Qwen3.8-Flash-Next的原生逐层嵌入子系统中。使用针对两个基准独立训练的检查点,BadEngram在HarmBench上实现了50.4%的ASR,在AdvBench上实现了60.0%的ASR,而休眠条件下的ASR分别保持在0.9%和0.0%。这些结果将原生门控记忆参数确定为模型的安全关键部分,其完整性无法从不变的主干网络推断出来。

英文摘要

To expand open-weight models' capacity without proportionally increasing computation, recent architectures incorporate gated parametric memory that retrieves learned values and injects them into intermediate representations. One representative design is Engram, which combines deterministic n-gram lookup with context-dependent gating over large learned memory tables. Despite these efficiency benefits, such modules create a distinct attack surface: their parameters can be modified independently of the backbone while directly shaping its computation. We introduce BadEngram, a post-training attack that exploits this surface to implant persistent, trigger-dependent behavior while leaving conventional backbone weights and the execution graph unchanged. We first establish the attack's feasibility and causally characterize its mechanism in a controlled Engram model, where BadEngram achieves 96.6% ASR on triggered inputs while limiting false activation on matched trigger-free inputs to 0.1% and preserving 99.6% clean accuracy. Replacing the retrieved memory values with their clean counterparts or closing the memory gates reduces ASR to at most 0.32%, confirming that the backdoor is expressed through the gated-memory pathway. We then test whether this vulnerability extends to production scale in Qwen3.8-Flash-Next's native Per-Layer Embedding subsystem. Using independently trained checkpoints for the two benchmarks, BadEngram achieves 47.8% ASR on HarmBench and 64.0% on AdvBench, while dormant-condition ASR remains 0.9% and 0.0%, respectively. These results identify native gated-memory parameters as a security-critical part of the model whose integrity cannot be inferred from an unchanged backbone.

补充信息

↑