孤立但暴露:基于持久性的大语言模型智能体内存提取攻击
Isolated but Exposed: Persistence-Based Memory Extraction Attack on LLM Agents
AI总结:
研究基于大语言模型的智能体内存提取攻击,提出SPORE攻击方法,通过解耦对抗命令与检索锚点、优化覆盖空间及持久化有效载荷,突破内存隔离,实现高提取率,揭示仅靠内存隔离不足,需重审工具端信任边界。
AI中文摘要:
基于大语言模型的智能体通过长期记忆(LTM)扩展了大语言模型,LTM会在多个会话中持久保存隐私敏感的用户数据。生产系统通过内存隔离来降低提取风险,将每个用户的LTM绑定到唯一标识符。我们发现工具接口是一个被忽视的攻击面。智能体经常将从LTM检索到的数据嵌入工具调用参数中,使恶意工具能够在不违反用户级隔离的情况下窃取私有内存。我们提出了SPORE,这是针对此威胁模型设计的首次提取攻击。SPORE通过在短期内存中持久保存命令并在工具响应中发出语义纯净的锚点,将对抗性命令与检索锚点解耦。恢复的检索精度能够在嵌入空间上进行几何覆盖优化,系统地将锚点导向未探索的内存区域。为了在超出工具调用限制的情况下持续提取,SPORE在内存中持久保存重新激活有效载荷,无需额外用户触发即可在会话内和跨会话自动恢复攻击。SPORE在无限制触发时实现了80.0%的创纪录提取率,仅用20次触发时为47.0%。在多用户部署中,攻击者可以将提取的记录与用户身份关联,实现有针对性的监视。这些结果表明仅内存隔离是不够的,需要重新审视智能体架构中工具端的信任边界。
英文摘要:
LLM-based agents extend large language models with long-term memory (LTM) that persists privacy-sensitive user data across sessions. Production systems mitigate extraction risks through memory isolation, binding each user's LTM to a unique identifier. This defense has blocked known attacks on shared storage, fostering the assumption that isolated LTM is secure. We identify the tool interface as an overlooked attack surface. Agents routinely embed LTM-retrieved data in tool invocation parameters, enabling a malicious tool to exfiltrate private memory without violating user-level isolation. Naive adaptations of user-side extraction techniques fail because the adversarial command's semantics interfere with retrieval precision, and platform-imposed tool-call limits constrain the extraction budget per trigger. We present SPORE, the first extraction attack designed for this threat model. SPORE decouples the adversarial command from retrieval anchors by persisting the command in short-term memory and emitting semantically pure anchors in tool responses. The restored retrieval precision enables a geometric coverage optimization over the embedding space that systematically steers anchors toward unexplored memory regions. To sustain extraction beyond tool-call limits, SPORE persists reactivation payloads in memory that automatically resume the attack within and across sessions without additional user triggers. SPORE achieves an 80.0% record extraction rate with unlimited triggers and 47.0% with only 20 triggers. In multi-user deployments, attackers can link extracted records to user identities, enabling targeted surveillance. These results demonstrate that memory isolation alone is insufficient and call for reexamining tool-side trust boundaries in agent architectures.