BMA:回链记忆攻击在LLM智能体中创建未授权控制路径
BMA: Backchain Memory Attacks Create Unauthorized Control Paths in LLM Agents
浏览论文内容
中文总结 AI 辅助
提出回链记忆攻击(BMA),通过逆向规划编辑低可信证据形成持久记忆,在干净任务中驱动未授权动作,实现18.8%路径认证成功率,并验证了基于来源的授权防御。
中文摘要 AI 辅助
持久记忆使LLM智能体能够复用先前经验,但同时也引入了新的安全边界:智能体可能记住的内容并不等同于其应当据以行动的内容。我们揭示了一条未授权控制路径,其中被编辑的低可信度证据被整合进持久记忆,在干净任务中被检索,并被用于驱动受保护的动作。关键在于,攻击者既不写入记忆也不篡改任务。我们提出了回链记忆攻击(BMA),这是一种灰盒、由LLM驱动的逆向规划攻击,从目标动作向后推理至可能触发该动作的记忆,再进一步推理至形成该记忆所需的证据编辑。BMA包含两个阶段:准备阶段使用可重置试验来定位首个失败环节并构建经验;执行阶段使用冻结的经验对候选编辑进行排序并提交,无需反馈。我们引入了路径认证攻击成功率(Path-CASR),以区分记忆介导的命中与偶然命中:注册的记忆必须形成、被检索、驱动目标行为,并通过匹配干预检查。在四个基底和三个决策主干上,BMA实现了18.8%的宏Path-CASR,而最强的访问匹配基线为13.4%。在BMA的行为命中中,60.3%通过了所有注册路径和干预检查,而基线为36.7%。冻结的BMA编辑在四个保留的整合策略上平均保留了其认证效果的78.0%。代表性的记忆侧控制留下了11.0%的Path-CASR,而基于来源的授权将其降至2.0%,同时保留了92.1%的合法动作成功率。
英文摘要
Persistent memory enables LLM agents to reuse prior experience, but creates a new security boundary: what an agent may remember is not what it should act on. We expose an unauthorized control path where edited low-trust evidence is consolidated into persistent memory, retrieved on a clean task, and used to drive a protected action. Crucially, the adversary neither writes memory nor alters the task. We introduce Backchain Memory Attack (BMA), a grey-box, LLM-driven inverse-planning attack that reasons backward from the target action to the memory that would trigger it, then to the evidence edit that would form it. BMA has two phases: preparation uses resettable trials to localize the first failed link and build experience; execution uses the frozen experience to rank and commit candidate edits without feedback. We introduce the Pathway-Certified Attack Success Rate (Path-CASR) to separate memory-mediated from coincidental hits: registered memory must form, be retrieved, drive the target behavior, and pass matched-intervention checks. Across four substrates and three decision backbones, BMA achieves 18.8% Macro Path-CASR, compared with 13.4% for the strongest access-matched baseline. Of BMA's behavioral hits, 60.3% pass all registered pathway and intervention checks versus 36.7% for the baseline. Frozen BMA edits retain 78.0% of their certified effect on average across four held-out consolidation policies. Representative memory-side controls leave 11.0% Path-CASR, whereas provenance-bound authorization reduces it to 2.0% while preserving 92.1% legitimate-action success.
发表机构
- School of Cyber Science and Technology, Harbin Institute of Technology(哈尔滨工业大学网络空间安全学院)
- SnT, University of Luxembourg(卢森堡大学科学与技术研究中心)
- Department of New Networks, Peng Cheng Laboratory(鹏城实验室新网络系)
机构由 AI 辅助整理,请以论文原文为准。