arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向大语言模型智能体的间接长期记忆投毒的可迁移端到端优化

Transferable End-to-End Optimization for Indirect Long-Term Memory Poisoning in LLM Agents

Chuanchao Zang, Jianing Wang, Wenyu Chen, Xiangtao Meng, Li Wang, Xinyu Gao, Zheng Li, Shanqing Guo

arXiv 2609.00523首次发表:更新:

发表机构

School of Cyber Science and Technology, Shandong University; State Key Laboratory of Cryptography and Digital Economy Security, Shandong University; Shandong Key Laboratory of Artificial Intelligence Security, Shandong University(山东大学网络空间安全学院; 山东大学密码与数字经济安全国家重点实验室; 山东大学人工智能安全山东省重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM智能体间接长期记忆投毒的阶段间耦合问题,提出端到端优化方法PipePoison,在多框架、多记忆机制下显著提升攻击利用率,且对未见过的受害者配置和防御措施仍有效。

AI 中文摘要

长期记忆可将不可信的外部内容转化为对大语言模型(LLM)智能体未来决策的持续影响,从而形成间接记忆投毒的威胁。一次成功的攻击必须经过包含记忆写入、检索和利用的多阶段流程。现有攻击方法大多依赖阶段内优化,即孤立地优化各个阶段,却忽略了阶段间的耦合关系。具体而言,这些阶段对同一份投毒内容有不同的要求,且每个阶段都基于其前序阶段的转换输出进行操作。因此,优化某一阶段可能会损害其他阶段的有效性,而上游转换可能会消除为下游阶段所做的改进。因此,间接记忆投毒应被视为一个端到端优化问题。基于这一见解,我们提出了PipePoison,该方法从本地影子系统收集细粒度的阶段反馈,使用链式结构的损失来识别和优化阻碍端到端成功的阶段瓶颈,并应用稳定性校准的阶段和配置权重来提高可迁移性。在三个智能体框架和四种记忆机制上,PipePoison将攻击利用率提高了19.1个百分点;即使在完全未见过的受害者配置上,它也比最强基线高出16个百分点,并且在八种代表性防御措施下仍然有效。

英文摘要

Long-term memory can turn untrusted external content into persistent influence over an LLM agent's future decisions, creating the threat of indirect memory poisoning. A successful attack must survive a multi-stage pipeline comprising memory writing, retrieval, and utilization. Existing attacks largely rely on intra-stage optimization, optimizing individual stages in isolation while overlooking inter-stage coupling. Specifically, these stages impose different requirements on the same poisoning content, and each stage operates on the transformed output of its predecessor. Consequently, optimizing one stage may undermine the effectiveness of other stages, while upstream transformations may erase improvements intended for downstream stages. Indirect memory poisoning should therefore be viewed as an end-to-end optimization problem. Based on this insight, we present \textsc{PipePoison}, which collects fine-grained stage feedback from local shadow systems, uses chain-structured losses to identify and optimize the stage bottlenecking end-to-end success, and applies stability-calibrated stage and configuration weights to improve transferability. Across three agent frameworks and four memory mechanisms, \textsc{PipePoison} improves attack utilization rate by 19.1 percentage points. Even on fully unseen victim configurations, it outperforms the strongest baseline by 16 percentage points and remains effective under eight representative defenses.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑