安全的代价:LLM智能体中记忆投毒防御的良性用例效用与Token开销
The Price of Safety: Benign-Case Utility and Token Overhead of Memory-Poisoning Defenses in LLM Agents
浏览论文内容
中文总结 AI 辅助
本研究测量LLM智能体记忆投毒防御在良性流量下的效用与Token开销,发现写入时防御无显著成本,而读取时重排序器导致准确率下降和误隔离,拦截位置决定代价。
中文摘要 AI 辅助
针对LLM智能体的记忆投毒防御通常以其阻止攻击的能力来评估。然而,它们处理的流量很少是敌对的。实施防御的成本在每次交互中都会付出,而其收益仅在很小比例的案例中显现。我们开发了一套测量设置,在不同条件下保持记忆后端、检索过程和评判器一致,仅改变防御本身。我们在五次对话中对每个条件测试三次,以区分防御的真实效果与管道运行中固有的噪声——即使在温度为零时,这种噪声仍然显著。在三种写入时防御(输入清洗、来源检查和基于LLM的异常检测)和一种读取时防御(重排序)中,在完全良性流量上测试,写入时防御没有显示出我们能够分辨的效用成本,95%置信区间约为±4.5分且包含零。重排序器则不同:它将核心准确率降低了4.4分(95% CI [-9.0,-0.05],bootstrap;McNemar p=0.064),这一结果在重复实验中依然存在,但处于我们分辨能力的边缘。其更清晰的成本是机械性的而非统计性的。在包含无攻击的对话中,重排序器在33.6%的判定项目上隔离了合法记忆,在单次对话中达到多达106次误隔离,Token开销为2.7%。将全部四种防御叠加并不会加剧这一成本:组合条件下的准确率损失更小,其置信区间包含零,表明写入时防御可能部分抵消了重排序器丢弃的内容。决定良性用例代价的似乎是防御在管道中的拦截位置,而非其是否使用LLM。
英文摘要
Memory-poisoning defenses for LLM agents are typically evaluated by their ability to prevent attacks. However, the traffic they process is rarely adversarial. The cost of implementing a defense is paid with each interaction, while its benefits are only seen in a small percentage of cases. We developed a measurement setup that keeps the memory backend, retrieval process, and judge consistent across different conditions, changing only the defense itself. We test each condition three times across five conversations to distinguish the defense's real effects from noise inherent in the pipeline's runs, which remains significant even at temperature zero. Across three write-time defenses (input sanitization, provenance checking, and LLM-based anomaly detection) and one read-time defense (reranking), tested on entirely benign traffic, the write-time defenses show no utility cost we can resolve, with 95% confidence intervals spanning roughly +/-4.5 points and including zero. The reranker is different: it lowers core accuracy by 4.4 points (95% CI [-9.0,-0.05], bootstrap; McNemar p=0.064), a result that survives replication but sits at the edge of our resolution. Its clearer cost is mechanical rather than statistical. On conversations containing no attack, the reranker quarantines legitimate memories on 33.6% of adjudicated items, reaching as many as 106 false quarantines in a single conversation, at 2.7% token overhead. Stacking all four defenses does not compound this cost: the combined condition's accuracy loss is smaller, and its confidence interval includes zero, suggesting the write-time defenses may partly offset what the reranker discards. Where a defense intercepts the pipeline, not whether it uses an LLM, appears to determine its benign-case price.