arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34132cs.AI

从攻击成功到攻击严重性:针对LLM智能体的反事实记忆攻击

From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents

  • Shanghai Academy of AI for Science (SAIS)(上海人工智能科学研究院)
  • Fudan University(复旦大学)
  • Monash University(莫纳什大学)

机构由 AI 辅助整理,请以论文原文为准。

Mingxi Zou, Langzhang Liang, Zhuo Wang, Yiyang Zhao, Lizhen Qu, Zenglin Xu

AI总结:

针对LLM智能体持久记忆攻击,提出以反事实记忆遗憾(CMR)衡量攻击严重性,并设计MemHarm方法,通过语义编辑和离线反馈优化,实现更高下游损失且保留成功率。

AI中文摘要:

随着LLM智能体越来越依赖持久记忆来实现长时程和个性化行为,它们能够在交互过程中保留并重用信息,但这也创造了一个持久的通道,使得恶意记忆写入能够影响未来的行为。持久记忆攻击通常通过是否成功来评估,然而成功的攻击可能留下具有显著不同下游后果的持久状态。我们将这种严重性作为一个独立的攻击设计目标进行研究,并用反事实记忆遗憾(CMR)将其形式化,即相对于干净记忆的预期下游损失的成对增加。我们引入了MemHarm,它预先声明了一类有限的稀疏、基于事实的语义编辑,通过正常的智能体记忆接口使用离线成对损失反馈来评估候选方案,并在此类内认证已解析的选择。与攻击成功优化相比,CMR引导的选择产生了显著更大的下游损失,同时保留了大部分成功率提升。在两个智能体基准测试和多种记忆设计中,MemHarm在相同支持度上达到了所有被评估的通用攻击中最高的CMR点估计。因子移除干预将该危害与所选的语义因子联系起来,原生智能体部署验证了写入到新进程的攻击路径。

英文摘要:

As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typically evaluated by whether they succeed, yet successful attacks can leave persistent states with substantially different downstream consequences. We study this severity as a distinct attack-design objective and formalize it with counterfactual memory regret (CMR), the paired increase in expected downstream loss relative to clean memory. We introduce MemHarm, which predeclares a finite class of sparse, grounded semantic edits, evaluates candidates through the normal agent memory interface using offline paired-loss feedback, and certifies resolved selections within that class. Compared with attack-success optimization, CMR-guided selection produces substantially larger downstream loss while retaining most of the success-rate gain. Across two agent benchmarks and diverse memory designs, MemHarm attains the highest CMR point estimates among the evaluated general attacks on identical support. Factor-removal interventions link this harm to the selected semantic factor, and native-agent deployments verify the write-to-fresh-process attack path.

↑