arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35576cs.AIcs.CLcs.CRcs.LG

共享载体AI病毒:跨LLM智能体的记忆跳跃攻击

Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents

Sidharth Pulipaka, Ansh Sharma, Stanislau Hlebik, Leonidas Raghav, Vyas Raina, Ivaxi Sheth, Mario Fritz

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示大型语言模型智能体通过共享持久化工件形成间接通信渠道,可导致自我传播的对抗性攻击,实验显示攻击能跨多个独立助手传播至八跳,影响60-80%的智能体。

中文摘要 AI 辅助

大型语言模型日益被部署为有状态助手,它们能在交互之间保留信息,并使用工具读取、修改和创建持久化工件。由于这些工件在用户之间共享,它们构成了原本相互独立的助手之间的间接通信渠道。我们研究了一种故障模式,其中该渠道能够实现自我传播的攻击。我们引入了工件介导的传播机制,即通过工件(如报告)引入的对抗性内容被存储在助手的持久记忆中,随后在新建的工件中重现,并被另一个稍后读取该工件的助手获取。我们在模拟独立运营的助手之间随时间进行工件交换的时间性人机智能体宇宙中评估了这一过程,测量攻击能否在连续交接中存活、能达到多少跳数以及传播范围有多广。我们发现攻击能够跨多个独立助手传播,并在扩展的交互序列中持续存在。在更大规模的模拟环境中,即使是GPT-5.6 Luna也表现出显著的传播,达到60-80%的智能体,传播链延伸至八跳。这些结果表明,持久化工件可以作为对抗性状态的持久载体,使攻击能够超越单个交互存活,并在孤立的助手之间传播。

英文摘要

Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent assistants. We study a failure mode in which this channel enables self-propagating attacks. We introduce artifact-mediated propagation, where adversarial content introduced through an artifact (e.g. a report), is stored in an assistant's persistent memory, reproduced in a subsequently created artifact, and acquired by another assistant that later reads it. We evaluate this process in temporal human-agent universes that model artifact exchange between independently operated assistants over time, measuring whether an attack survives successive hand-offs, how many hops it reaches, and how broadly it spreads. We find that attacks can propagate across multiple independent assistants and persist over extended interaction sequences. In larger simulated environments, even GPT-5.6 Luna exhibits substantial spread, reaching 60-80% of agents with propagation chains extending to eight hops. These results show that persistent artifacts can act as durable carriers of adversarial state, allowing attacks to outlive individual interactions and spread across isolated assistants.

发表机构

  • University of Cambridge(剑桥大学)
  • APTA AI
  • CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑