发表机构
University of California, Irvine(加州大学尔湾分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对多模态人工智能智能体长期记忆的黑盒视觉攻击,提出Lucid框架,通过制作扰动实现记忆中毒和注入两种攻击模式,在多种对话领域和架构上评估,揭示了多模态记忆管道的结构漏洞。
AI 中文摘要
多模态人工智能智能体越来越依赖持久的长期记忆来根据过去的视觉和文本情节进行生成。我们表明,对视觉数据的无条件信任会造成一个关键漏洞。我们提出了Lucid,一个黑盒对抗框架,它在严格的图像受限威胁模型下破坏多模态记忆管道,无需访问目标多模态语言模型、目标检索编码器或文本通道。Lucid精心制作难以察觉的扰动,基于历史上下文的可用性实现两种不同的失败模式:(1)记忆中毒,一种上下文攻击,对抗性图像取代良性图像,其内容由先前的文本上下文强化,可靠地破坏视觉回忆并引导智能体走向攻击者选择的叙述;(2)记忆注入,一种上下文外攻击,对抗性图像在没有先前文本基础的对话轮次中取代良性图像,导致智能体生成受攻击者影响的响应,且没有来自记忆的校正信号。我们在各种对话领域和五种黑盒记忆架构上评估了Lucid,包括图结构、由语言模型总结的和商业部署的系统。Lucid在中毒攻击上的ASR达到61.6%,在注入攻击上的ASR达到58.4%,揭示了多模态记忆管道中的结构漏洞。
英文摘要
Multimodal AI agents increasingly rely on persistent long-term memory to ground generation in past visual and textual episodes. We show that unconditional trust in visual data creates a critical vulnerability. We propose Lucid, a black-box adversarial framework that compromises multimodal memory pipelines under a strictly image-bounded threat model, requiring no access to the target MLLM, target retrieval encoder, or the text channel. Lucid crafts imperceptible perturbations to enable two distinct failure modes based on the availability of historical context: (1) Memory poisoning, an in-context attack where the adversarial image replaces a benign one whose content is reinforced by prior textual context, reliably corrupting visual recall and steering the agent toward attacker-chosen narratives; (2) Memory injection, an out-of-context attack where the adversarial image replaces a benign one in a conversation turn devoid of prior textual grounding, causing the agent to generate attacker-influenced responses with no corrective signal from memory. We evaluate Lucid across various conversation domains and five black-box memory architectures, including graph-structured, LLM-summarized, and commercially deployed systems. Lucid achieves 61.6% ASR on poisoning and 58.4% ASR on injection, exposing a structural vulnerability in multimodal memory pipelines.
Comments34 pages, 5 figures, 15 tables