受攻击下的效用:智能体记忆投毒及内容筛选与来源排序的局限性
Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking
查看机构详情
- Quantify Labs Ltd(Quantify实验室有限公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对智能体记忆投毒问题,测试了内容筛选和来源排序的防御效果,发现仅内容筛选无法抵御投毒,来源加权检索也存在局限,主张采用有限占用约束并发布了相关资源。
中文摘要 AI 辅助
持久记忆使虚假信息具有持久性:虚假陈述一旦被存储,就会在与其匹配的未来会话中被检索出来。我们使用单次生成的措辞直白的虚假断言来衡量这种失效模式的代价,这些虚假断言没有指令、触发词或检索器优化。对LongMemEval语料库的1.2%进行投毒,会使准确率从0.850降至0.300。一个四阶段写入时筛选管道在间接提示注入上达到0.832的召回率,同时标记1.5%带有触发词的良性文本,但对360个投毒记忆的拒绝数为0。我们认为这暴露了仅基于内容筛选的局限性:区分虚假断言与真实断言通常需要文本之外的外部依据。随后我们评估了来源加权检索。默认权重与无防御措施在统计上无差异(p=0.80),而更强的权重仅在排除不可信内容时才能恢复效用。在不可信内容大多为良性的混合来源语料库中,准确率从0.3167升至0.7000;当承载答案的证据本身来自不可信来源时,证据召回率降至0,准确率降至0.0417。在测得的相似性机制下,加性来源项没有可用的设置:足以抵抗查询型投毒的权重也会强到抑制合法的不可信证据。因此我们主张在检索时采用有限占用约束而非加性来源惩罚,并发布了测试工具、语料库及汇总运行报告。
英文摘要
Persistent memory makes false information durable: once a false statement is stored, it can be retrieved into future sessions that match it. We measure the cost of this failure mode using plainly worded false assertions generated in a single pass, with no instruction, trigger, or retriever optimization. Poisoning 1.2% of a LongMemEval corpus reduces accuracy from 0.850 to 0.300. A four-stage write-time screening pipeline that reaches 0.832 recall on indirect prompt injection while flagging 1.5% of trigger-word-laden benign text rejects 0 of 360 poisoned memories. We argue this exposes a boundary of content-only screening: distinguishing a false assertion from a true one generally requires external grounding beyond the text itself. We then evaluate provenance-weighted retrieval. The shipped weight is statistically indistinguishable from no defense (p=0.80), while a stronger weight recovers utility only by excluding untrusted content. In a mixed-provenance corpus where untrusted content is mostly benign, accuracy rises from 0.3167 to 0.7000; when the answer-bearing evidence itself arrives untrusted, evidence recall falls to zero and accuracy to 0.0417. Under the measured similarity regime, the additive provenance term has no usable setting: a weight strong enough to resist query-shaped poison is also strong enough to suppress legitimate untrusted evidence. We therefore argue for bounded occupancy constraints at retrieval rather than additive provenance penalties, and release the harnesses, corpora, and aggregate run reports.