arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09027cs.CRcs.AIcs.CV

视觉记忆攻击可以通过KV缓存持续存在

Visual Memory Attacks Can Persist Through The KV Cache

David Dobre, Leo Schwinn, Gauthier Gidel, Spandana Gella, Perouz Taslakian, Pierre-André Noël

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出P-VMI攻击,证明对抗性图像可通过KV缓存持续影响模型,即使从上下文中移除后仍有效,在Qwen3-VL-8B上达到90%成功率。

中文摘要 AI 辅助

现代语言模型系统在包含不可信文本和图像的超长上下文中自主运行。对抗性输入能否在从上下文中移除后继续引导模型?我们证明,攻击可以被训练为通过后续令牌的键/值(KV)缓存持续存在,使得对抗性影响在无法直接访问其来源后仍然有效。我们考虑视觉记忆注入(VMI;Schlarmann和Hein,2026)攻击设置,其中上下文中保留的对抗性图像植入隐藏后门:模型表现正常,直到选定的触发器引发攻击者选择的响应。我们首先通过优化的软提示在此设置中展示持续性,这些软提示在我们将提示从注意力中屏蔽后仍然有效。然后,我们引入持久性视觉记忆注入(P-VMI),该方法优化图像以在从注意力中屏蔽后保持这种对抗性行为。这些攻击在比优化期间使用的对话更长的对话中持续存在。在Qwen3-VL-8B-Instruct上,P-VMI在其最强配置中达到约90%的目标成功率,并且在更严格的移除设置下(仅在第一轮暴露图像)仍然有效。缓存交换消融将持久性影响定位到KV缓存。最后,我们表明这些攻击可以被训练为在保留由同一模型生成的摘要的KV缓存的压缩中存活,证明对抗性行为可以在缓存状态中持续存在,而无需持续访问其来源。

英文摘要

Modern language model systems operate autonomously over increasingly long contexts containing untrusted text and images. Can an adversarial input continue to steer a model even after that input is removed from its context? We show that attacks can be trained to persist through the key/value (KV) cache of subsequent tokens, allowing adversarial influence to outlive direct access to its source.We consider the Visual Memory Injection (VMI; Schlarmann and Hein, 2026) attack setting, in which an adversarial image that stays in the context plants a hidden backdoor: the model behaves normally until a chosen trigger elicits an attacker-chosen response. We first demonstrate persistence in this setting with optimized soft prompts, which remain effective after we mask the prompt from attention. We then introduce Persistent Visual Memory Injection (P-VMI), which optimizes images to preserve this adversarial behaviour after they are masked from attention. These attacks persist over conversations substantially longer than those used during optimization. On Qwen3-VL-8B-Instruct, P-VMI achieves up to approximately $90\%$ target success in its strongest configuration and remains effective under a stricter removal setting that exposes the image only on the first turn. A cache-swap ablation localizes the persistent influence to the KV cache. Finally, we show that these attacks can be trained to survive compaction that retains the KV cache of a summary generated by the same model, demonstrating that adversarial behaviour can persist in cached state without continued access to its source.

发表机构

  • ServiceNow AI Research(ServiceNow人工智能研究院)
  • Université de Montréal(蒙特利尔大学)
  • Mila(米拉)
  • Technical University of Munich(慕尼黑工业大学)
  • McGill University(麦吉尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑