arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14651cs.CRcs.AI

MemPoison:揭示大语言模型智能体中持久内存威胁和结构盲点

MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents

Jifeng Gao, Kang Xia, Yi Zhang, Xiaobin Hong, Mingkai Lin, Xingshen Wei, Wenzhong Li, Sanglu Lu

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型智能体持久内存安全漏洞,提出MemPoison框架,含1227个案例,涵盖多种攻击类型等并评估多个模型系列。引入分类法,揭示防御边界及盲点,主张转向自适应、上下文敏感的内存防御策略。

中文摘要 AI 辅助

持久外部内存增强了智能体的连续性,但引入了持久的安全漏洞:对抗性内容可通过标准交互通道注入,跨轮保留并扭曲下游行为。为应对这一挑战,我们提出MemPoison,一个全面的基准测试和分析框架,包含1227个经过人工验证的案例,涵盖四种攻击类型、三种注入通道和三种代表性内存基板,在七个开放权重和三个封闭权重模型系列上进行评估。我们引入了三层分类法:(L1)直接单记录损坏、(L2)组合多记录损坏和(L3)上下文触发的休眠损坏。评估揭示了一个明显的防御边界:虽然基线写入时防御(如一致性检查)大幅抑制直接L1攻击,但无法可靠抑制L2和L3攻击。通过机制影响分解(MID),我们展示了写入时防御中的结构盲点,这些盲点允许看似良性的记录通过联合检索组合或触发条件激活后来变得有害。我们的发现主张从静态过滤转向自适应、上下文敏感的内存防御策略。

英文摘要

Persistent external memory enhances agent continuity but introduces persistent security vulnerabilities: adversarial content can be injected via standard interaction channels, retained across turns, and later distort downstream behavior. To address this challenge, we propose MemPoison, a comprehensive benchmark and analysis framework featuring 1227 hand-validated cases across four attack types, three injection channels, and three representative memory substrates, evaluated on seven open-weight and three closed-weight model families. We introduce a three-tier taxonomy: (L1) direct single-record corruption, (L2) compositional multi-record corruption and (L3) context-triggered dormant corruption. Our evaluations reveal a distinct defense frontier: while baseline write-time defenses, such as consistency checks, substantially suppress direct L1 attacks, they fail to reliably suppress L2 and L3 attacks. Through mechanistic influence decomposition (MID), we demonstrate structural blind spots in write-time defenses, which admit seemingly benign records that later become harmful through joint retrieval composition or trigger-conditioned activation. Our findings advocate for shifting from static filtering to adaptive, context-sensitive memory defense strategies.

发表机构

  • State Key Laboratory for Novel Software Technology, Nanjing University, China(南京大学新型软件技术国家重点实验室)
  • NARI Group Corporation/State Grid Electric Power Research Institute, China(国网电力研究院/国电电力集团)

机构由 AI 辅助整理,请以论文原文为准。

↑