发表机构
Institute of Cyberspace Security, Zhejiang University of Technology; Binjiang Institute of Artificial Intelligence, ZJUT; College of Information Engineering, Zhejiang University of Technology(浙江工业大学网络空间安全研究院; 浙江工业大学滨江人工智能研究院; 浙江工业大学信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出MemSecBench基准测试,通过24种配置的实验发现不同内存系统栈在智能体内存中毒的持久性、后果及选择性修复上存在显著差异,为智能体内存安全评估提供了工具。
AI 中文摘要
内存系统使智能体能够保留和复用过往交互的信息,但也可能让恶意内容持续存在。攻击者精心构造的恶意指令可能被存储在长期内存中,在很久之后被召回并悄然影响实际行动。近期的基准测试越来越多地研究智能体内存安全,但很少在不同内存后端的对比中,追踪同一恶意语义在持久性、下游后果及选择性修复中的表现。为填补这一空白,我们提出MemSecBench,这是一个基于任务的智能体内存系统生命周期安全基准测试,包含来自代码、科学、日常生活及办公场景共48种现实情境的310个案例。每个案例在隔离运行时中遵循受控的写入-执行-遗忘协议,采用由智能体 harness、内存后端和大语言模型(LLM)后端定义的精确智能体配置。基于证据的裁决结合确定性写入检查、特定检查点的评判模型评估,以及跨7个生命周期检查点的程序化闸门。实验设计涵盖24种配置矩阵,包含2种智能体 harness、4种内存后端和3种LLM后端。在所有24种配置中,84.2%的案例存在恶意内存持续,完整的写入-执行链成功率为50.3%;在成功中毒的案例中,59.6%完成完整的执行链,56.1%实现选择性修复;在匹配Native配置下,端到端攻击成功的最大绝对差异为16.1个百分点,选择性修复的最大绝对差异为41.3个百分点。这些描述性对比表明,被评估的内存系统栈在生命周期安全方面存在差异,既体现在恶意内存的传播上,也体现在内存中毒成功后的选择性修复上。
英文摘要
Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A malicious instruction crafted by an attacker may be stored in long-term memory, recalled much later, and quietly shape a real action. Recent benchmarks increasingly examine agent memory security, yet few trace the same malicious semantics across persistence, downstream consequences, and selective repair under diverse memory-backend comparisons. To address this gap, we introduce MemSecBench, a task-grounded benchmark for the lifecycle security of agent memory systems. It contains 310 cases drawn from 48 realistic contexts across code and science, daily life, and office work. Each case follows a controlled Write--Execute--Forget protocol in an isolated runtime under an exact agent configuration, defined by an agent harness, a memory backend, and an LLM backend. Evidence-based adjudication combines a deterministic write check, checkpoint-specific judge-model evaluations, and programmatic gates across seven lifecycle checkpoints. The experimental design spans a 24-configuration matrix of two agent harnesses, four memory backends, and three LLM backends. Across all 24 configurations, malicious memory persists in 84.2% of all cases, and the full Write--Execute chain succeeds in 50.3%. Among successfully poisoned cases, 59.6% complete the full Execute chain, while 56.1% achieve selective repair.Compared with matched Native configurations, the largest absolute differences are 16.1 percentage points for end-to-end attack success and 41.3 percentage points for selective repair. These descriptive contrasts indicate that the evaluated memory system stacks differ in lifecycle security, both in the propagation of malicious memory and in selective repair after successful memory poisoning.