arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当恶意指令持续存在:针对基于 Harness 的智能体的持久内存投毒攻击

When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents

Shuhuai Huang, Jingfeng Zhang, Hong Jia

arXiv 2609.13889首次发表:更新:

发表机构

University of Auckland; Fudan University(奥克兰大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出针对基于 Harness 的智能体的持久内存投毒攻击 PMPA,通过诱导写入恶意指令实现跨会话隐私泄露,在 OpenClaw 和 Claude Code 上验证了高成功率,并评估了防御的局限性。

AI 中文摘要

Harness 设计通过整合内存、工具使用和运行时控制,彻底改变了基于 LLM 的智能体的开发方式。然而,这种设计也引入了安全和隐私风险,因为来自外部来源的恶意指令可能被写入持久内存,并在多个会话中持续存在。为了研究这一风险,我们提出了 PMPA,一种针对基于 Harness 的智能体的持久内存投毒攻击。PMPA 将恶意指令嵌入到良性的外部来源中,并诱导受害智能体在不直接访问智能体框架的情况下将其写入持久内存。一旦存储,被投毒的内存可以在后续会话中被检索,触发额外的恶意操作并导致隐私泄露。我们在 OpenClaw 和 Claude Code 上,针对不同的骨干 LLM、输入模态和触发场景评估了 PMPA。在所有设置中,PMPA 在 OpenClaw 上实现了平均注入成功率(ISR)和跨会话攻击成功率(C-ASR)分别为 73.7% 和 55.5%,在 Claude Code 上分别为 66.9% 和 81.7%,同时在这两个系统上保持了良性任务的性能。我们进一步评估了一种针对性的提示级防御,发现它在许多设置中可以减少内存注入,但一旦持久内存已被投毒,其提供的保护有限。

英文摘要

Harness design has transformed the development of LLM-based agents by integrating memory, tool use, and runtime control. However, this design also introduces security and privacy risks because malicious instructions from external sources may be written into persistent memory and persist across sessions. To study this risk, we propose PMPA, a Persistent Memory Poisoning Attack against harness-based agents. PMPA embeds malicious instructions into benign external sources and induces the victim agent to write them into persistent memory without directly accessing to the agent framework. Once stored, the poisoned memory can be retrieved in later sessions, triggering additional malicious actions and causing privacy leakage. We evaluate PMPA on OpenClaw and Claude Code across different backbone LLMs, input modalities, and trigger scenarios. Across all settings, PMPA achieves average Injection Success Rate (ISR) and Cross-session Attack Success Rate (C-ASR) of 73.7%/ 55.5% on OpenClaw and 66.9%/ 81.7% on Claude Code, while preserving benign task performance on both systems. We further evaluate a targeted prompt-level defense and find that it can reduce memory injection in many settings, but provides limited protection once the persistent memory has been poisoned.

Comments14 pages, 1 figures. Code: https://github.com/hsh754/PMPA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑