arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型智能体中的内存来源洗白:用于持久内存的非放大防火墙

Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory

Jinghan Xu, Yiyong Xiao, Wanru Shao, Hankai Liu, Xinjin Li

arXiv 2607.29167首次发表:更新:

AI 中文总结

针对LLM智能体内存来源洗白的安全问题,提出PPMF防火墙,通过匹配行动风险与记忆权限,成功阻止高风险未授权行动,保障智能体内存安全。

AI 中文摘要

长期记忆使大语言模型(LLM)智能体重用先前的偏好和工作流程,但也会将不可信的观测结果转化为持久的行动上下文。我们识别出内存来源洗白:在基于LLM的内存整合过程中,外部观测结果可能被重写为看似用户历史或工作流支持,保留行动触发条件的同时抹去应限制其权限的低信任来源。现有的提示过滤器、内容消毒器和工具守卫无法在有损内存整合后执行来源权限的非放大。我们将此边界形式化,并实例化为来源保留内存防火墙(PPMF),这是一种轻量级内存中间件,可保留平台维护的来源,并通过将行动风险与相关记忆的权限匹配来授权工具调用。在我们基于模式的固定风险策略评估中,易受攻击的整合记忆达到1.000的攻击成功率(ASR);借助完整的平台维护来源、确认信息和风险标签,未评估的未授权高风险行动无法通过PPMF网关,而确认的良性行动和目标低风险内存使用仍可执行。

英文摘要

Long-term memory lets large language model(LLM) agents reuse prior preferences and work flows, but it also turns untrusted observations into persistent action context. We identify memory provenance laundering: during LLM-based memory consolidation, an external observation may be rewritten as apparent user history or workflow support, preserving an action trigger while erasing the low-trust source that should limit its authority. Existing prompt filters, content sanitizers, and tool guards do not enforce source-authority non-amplification after lossy memory consolidation. We formalize this boundary and instantiate it as Provenance-Preserving Memory Fire wall (PPMF), a lightweight memory middleware that preserves platform-maintained provenance and authorizes tool calls by matching action risk to the authority of action-relevant memories. In our schema-grounded evaluation with fixed risk policies, vulnerable consolidated memories reach up to 1.000 attack success rate(ASR); with intact platform-maintained provenance, confirmation, and risk labels, no evaluated unauthorized high-risk action passes the PPMF gate while confirmed benign actions and targeted low-risk memory use remain executable.

CommentsEMNLP2026 submitted

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑