AI 中文总结
针对OpenClaw,提出MemCollusion框架,采用Salami策略生成合谋式内存投毒攻击,在48种场景下平均内存保存率81.3%、攻击成功率75.0%,可应对良性内存稀释与内存级防御。
AI 中文摘要
长期记忆使大语言模型(LLM)智能体能够在不同会话间保留有用信息,但也为攻击者提供了攻击面,攻击者可通过该攻击面投毒智能体的持久记忆以操纵其行为。现有内存投毒攻击主要依赖单独的恶意记录,却忽略了一种组合式威胁:多个看似良性的记忆可能共同诱导不安全行为。本文提出MemCollusion,一种用于构建合谋式内存投毒攻击的自动化红队框架。MemCollusion采用Salami策略——即将攻击目标拆分为多个单独无害的小片段的策略——来生成单独看似良性但整体有害的记忆片段。它利用四个设计约束、五种基于理论的策略及一个微调后的生成器来构建记忆联盟。为在现实的跨会话场景中评估合谋式内存投毒,我们开发了MoltLab,这是对Moltbook的受控研究复现,其中精心设计的平台内容必须先被观察并提炼为持久记忆,才能在后续会话中影响智能体行为。我们在OpenClaw上针对48种场景、使用两个主干模型评估MemCollusion,在最强内存保存设置下,MemCollusion的平均内存保存率达81.3%,攻击成功率达75.0%,且在良性内存稀释和内存级防御下仍保持有效。
英文摘要
Long-term memory enables LLM agents to retain useful information across sessions, but also creates an attack surface through which adversaries may poison an agent's persistent memory to steer its behavior. Existing memory poisoning attacks mainly rely on individually malicious records, overlooking a compositional threat: multiple benign-looking memories may jointly induce unsafe behavior. In this paper, we introduce MemCollusion, an automated red-teaming framework for constructing collusive memory poisoning attacks. MemCollusion applies salami tactics---a strategy that slices an adversarial objective into small, individually innocuous pieces---to generate memory fragments that are individually benign looking but collectively harmful. It constructs memory coalitions using four design constraints, five theory-informed strategies, and a fine-tuned generator. To assess collusive memory poisoning in a realistic cross-session setting, we develop MoltLab, a controlled research reproduction of Moltbook, in which crafted platform content must first be observed and distilled into persistent memory before influencing the agent's behavior in a separate session. We evaluate MemCollusion on OpenClaw using two backbone models across 48 scenarios. Under the strongest memory-saving setting, MemCollusion achieves an average Memory Save Rate of 81.3% and an Attack Success Rate of 75.0%, and remains effective under both benign memory dilution and memory-level defenses.