arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14611cs.CRcs.AIcs.MA

糟糕的记忆:评估智能体系统中内存引发的提示注入风险

Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems

  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

Soham Gadgil, David Alexander, Sai Sunku, Franziska Roesner

AI总结:

研究基于内存的智能体系统中提示注入攻击,利用沙盒合成工作区评估两个系统四个模型,发现虽难用外部内容重写内存文件,但已植入的有效载荷可攻击当前及未来会话,揭示持久内存改变威胁模型并推动相关防御研究。

AI中文摘要:

一类不断发展的智能体系统通过内存文件、行为偏好和知识库在会话间维持持久状态。这虽使智能体更有用且能自我改进,但也为提示注入创造了新的攻击面,恶意指令可嵌入持久文件并影响未来行为。本文利用沙盒合成工作区研究基于内存的智能体系统中的提示注入攻击。我们评估了Anthropic Claude Code和OpenAI Codex这两个智能体系统的四个模型:Claude Haiku 4.5、Claude Opus 4.7、GPT - 5.2和GPT - 5.5。结果表明,虽难以用不可信外部内容使智能体重写自身内存文件,但已植入文件的有效载荷可成功攻击当前及未来会话。攻击成功率和有效载荷持久性在不同系统、模型、对抗目标和多会话攻击序列中差异很大。这些发现表明持久内存改变了提示注入的威胁模型,并促使开发在不消除智能体有益适应性的情况下保护内存更新的防御措施。

英文摘要:

A growing class of agentic systems maintain persistent state across sessions through memory files, behavioral preferences, and knowledge bases. While this makes agents more useful and self-improving, it also creates a new attack surface for prompt injections in which malicious instructions can be embedded within persistent files and influence future behavior. In this work, we study prompt injection attacks in memory-based agentic systems using a sandboxed synthetic workspace. We evaluate two agentic systems, Anthropic Claude Code and OpenAI Codex, across four models: Claude Haiku 4.5, Claude Opus 4.7, GPT-5.2, and GPT-5.5. Our results show that although it is difficult to make an agent overwrite its own memory files using untrusted external content, payloads already planted in those files can successfully attack current and future sessions. Attack success and payload persistence vary substantially across systems, models, adversarial goals, and multi-session attack sequences. These findings show that persistent memory changes the threat model for prompt injection and motivate defenses that protect memory updates without removing useful agent adaptation.

补充信息

↑