AI 中文总结
针对LLM智能体记忆注入攻击,提出轻量防御框架MIND,通过意图感知信息瓶颈过滤冗余信息并识别恶意记忆,在ReAct-StrategyQA上大幅降低攻击成功率且保留任务性能与效率。
AI 中文摘要
基于大语言模型(LLM)的记忆增强智能体易受记忆注入攻击:智能体可能检索到攻击者投毒的记忆,这会使其行为偏离初始用户意图,最终导致任务失败。然而,现有防御机制要么带来高计算开销,要么在多轮对话上下文中存在信息冗余。为应对这些挑战,我们提出Memory Intent-Aware Neural Denoising(MIND),一种用于记忆注入攻击的轻量防御框架。我们的初步分析表明,良性和投毒轨迹在初始用户意图与后续行为之间表现出可区分的关系。基于这一观察,MIND采用意图感知信息瓶颈(IB)从初始意图和轮次级行为中提取紧凑的意图-行为表示。该IB保留与意图相关的跨轮攻击信号,同时过滤与任务无关的重复信息,一个轻量检测器从生成的表示中识别恶意记忆。因此,MIND在缓解多轮上下文信息冗余的同时,避免了重复LLM审计的开销。大量实验表明,MIND在降低攻击成功率的同时保留了任务准确率和推理效率。值得注意的是,在ReAct-StrategyQA上,MIND将平均ASR-r和ASR-a分别降低55.4%和55.3%,同时在平均准确率和延迟上与未防御的智能体相当。
英文摘要
Memory-augmented LLM-based agents are vulnerable to memory injection attacks: Agents may retrieve poisoned memory from attackers, which diverts their behavior from initial user intent and finally causes task failure. However, existing defense mechanisms either incur high computational cost or suffer from information redundancy in multi-turn contexts. To address these challenges, we propose Memory Intent-Aware Neural Denoising(MIND), a lightweight defense framework for memory injection attack. Our preliminary analysis reveals that benign and poisoned trajectories exhibit distinguishable relationships between the initial user intent and subsequent behavior. Building on this observation, MIND employs an intent-aware Information Bottleneck(IB) to extract compact intent--behavior representations from the initial intent and turn-level behavior. The IB preserves intent-relevant cross-turn attack signals while filtering task-irrelevant and repetitive information, and a lightweight detector identifies malicious memories from the resulting representations. As such, MIND mitigates information redundancy in multi-turn contexts while avoiding the overhead of repeated LLM auditing. Extensive experiments show that MIND reduces attack success rates while preserving task accuracy and inference efficiency. Notably, on ReAct-StrategyQA, MIND reduces mean ASR-r and ASR-a by 55.4% and 55.3%, respectively, while matching the undefended agent in average accuracy and latency.