InjecMEM:针对大语言模型智能体记忆系统的内存注入攻击
InjecMEM: Memory Injection Attack on LLM Agent Memory Systems
浏览论文内容
中文总结 AI 辅助
本研究提出InjecMEM内存注入攻击,无需访问LLM智能体内存存储,通过特定锚点与对抗性命令引导后续查询输出,在多模型上验证了其有效性,为智能体内存安全研究提供了可复现框架。
中文摘要 AI 辅助
内存正成为部署后的大语言模型(LLM)智能体中提供持久个性化与连续性的默认子系统,这自然引发一个问题:内存系统是否会给智能体引入新的漏洞?为此,我们提出InjecMEM,一种新颖的内存注入攻击范式,仅需一次交互(无需对内存存储进行读取/编辑访问)即可将后续相关查询的响应引导至指定输出。受内存系统“先检索后生成”机制的指导,我们采用与检索器无关的锚点和对抗性命令来构建注入内容:锚点包含高召回率的主题线索,确保下游检索始终将该记录与目标主题关联;命令是一段短序列,经优化后可在不确定的融合上下文、可变位置及长提示下保持有效,确保其在被检索后能可靠引导输出。我们通过基于梯度的坐标搜索学习该命令,在合成提示模板与插入位置上取平均,并将其扩展至跨主干模型的联合优化以研究迁移性。在多个内存系统与主干模型上的评估显示,InjecMEM可实现可靠的主题条件检索与目标生成,在内存漂移下仍保持有效,且不会影响非目标查询。我们的结果强调了加固内存系统的必要性,并提供了用于研究智能体记忆的可复现框架。
英文摘要
Memory is becoming a default subsystem in deployed LLM agents to provide persistent personalization and continuity. This naturally prompts a question: will memory system introduce new vulnerabilities into agents? Thus we propose InjecMEM, a novel memory injection attack paradigm that requires only a single interaction (no read/edit access to memory store) to steer later responses of related queries toward a pre-specified output. Guided by the retrieval-then-generate mechanism of memory systems, we craft the injection with a retriever-agnostic anchor and an adversarial command. The anchor contains high-recall topical cues so that downstream retrieval consistently associates the record with the target topic. The command is a short sequence optimized to remain effective under uncertain fused contexts, variable placements, and long prompts so that it reliably steers outputs once retrieved. We learn the command via gradient-based coordinate search, averaging over synthetic prompt templates and insertion positions, and extend it to joint optimization across backbones to study transfer. Evaluated across multiple memory systems and backbone models, InjecMEM achieves reliable topic-conditioned retrieval and targeted generation, remains effective under memory drift, and leaves non-target queries unaffected. Our results underscore the need to harden memory systems and provide a reproducible framework for studying agent memory.
发表机构
- Institute of Image Processing and Pattern Recognition, Shanghai Jiao Tong University(上海交通大学图像处理与模式识别研究所)
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。