发表机构
Vulcan Research, AIFT(伏尔肯研究所,AIFT)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文实证研究五个代理记忆系统,发现软撤销标记在检索时未被强制执行,导致代理采取不安全行动,并提出位于代理与记忆后端之间的守卫来扣留冲突记录。
AI 中文摘要
长期运行的语言模型代理依赖于持久记忆。许多代理记忆系统通过软撤销来保留历史:被矛盾的事实标记为无效并保留而非删除。然而,该标记在检索时是否被执行尚未得到检验。在本文中,我们测量了五个此类系统:我们向每个系统加载一个已撤销的策略及其替代策略,追踪在检索时是否返回已撤销的事实,以及代理是否在九个策略场景和九个模型下据此行动,并在六种防御条件下对每次试验进行评分。我们发现,没有系统默认执行撤销:只要撤销标签对检索层可见,已撤销的事实就会被返回,其排名高于替代策略,并引导代理采取不安全行动。基于这些发现,我们开发了一个守卫,它位于代理与任何记忆后端之间,并扣留已撤销或与其替代策略冲突的记录。
英文摘要
Long-running language-model agents depend on persistent memory. Many agent-memory systems preserve history through soft revocation: a contradicted fact is marked invalid and retained rather than deleted. However, whether that mark is enforced at retrieval time is unexamined. In this paper, we measure five such systems: we load each with a revoked policy and its replacement, track whether the revoked fact is returned at retrieval and whether the agent then acts on it across nine policy scenarios and nine models, and score every trial under six defense conditions. We find that no system enforces revocation by default: the revoked fact is returned wherever the revocation label is visible to the retrieval layer, outranks its replacement, and leads agents to the unsafe action. Based on these findings, we develop a guard that sits between the agent and any memory backend and withholds records that are revoked or conflict with their replacement.
CommentsCode is available at https://github.com/VulcanLab/Memory-Rebirth-Attack