arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相似性并非有效性:防御LLM语义缓存投毒

Similarity Is Not Validity: Defending LLM Semantic Caches Against Poisoning

Zihan Zhang, Shuangjie Yao, Zesen Liu, Zhixiang Zhang, Wai Ip Lai, Dung Hiu Hilton Yeung, Chun Kit Zhang, Fuchen Ma, Yuanyuan Yuan, Yu Jiang, Dongdong She

arXiv 2609.35908首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; Tsinghua University(香港科技大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM语义缓存投毒攻击,提出基于删除增益和答案检查的防御方法,从缓存键原始文本恢复信息,在5%误报率下阻断82.0%-98.2%投毒条目,开销可忽略。

AI 中文摘要

语义缓存通过复用先前为语义相似查询生成的答案来降低LLM服务成本。然而,检索仅基于输入查询与缓存查询之间的嵌入相似性。这种设计导致了缓存投毒:攻击者可以在与良性请求具有高余弦相似度的查询下缓存恶意响应。该漏洞源于检索相似性与答案有效性之间的差距。从信息瓶颈的角度来看,查询嵌入可能会丢失区分有效与无效缓存命中所需的信息,这限制了任何仅使用这些嵌入的匹配算法。我们提出了一种新颖的防御方法,从缓存键的原始文本中恢复这一必要信息。在各种投毒攻击中,对抗性查询共享一种重写-残差结构:它们将目标查询的重写与残差内容配对。重写保持了高相似性,而残差则引发恶意响应。删除残差会使剩余的重写与输入查询更加相似。我们利用这一结构,通过删除增益来搜索缓存查询的缩短变体以获取相似性增益,并使用答案检查来测试被删除的文本是否对存储的答案有贡献。我们证明了当删除使文本足够接近重写时,删除增益保持为正,并且我们使用滑动窗口搜索此类删除。在三种投毒攻击类别中,我们的防御在5%的误报率下阻止了82.0%至98.2%的投毒条目,且服务开销可忽略不计。

英文摘要

Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries. However, retrieval is based solely on embedding similarity between the incoming query and cached queries. This design enables cache poisoning: an attacker can cache a malicious response under a query with high cosine similarity to benign requests. The vulnerability stems from a gap between retrieval similarity and answer validity. From an information-bottleneck perspective, query embeddings can lose information needed to distinguish valid from invalid cache hits, which limits any matching algorithm that uses only these embeddings. We propose a novel defense that recovers this necessary information from the raw text of the cache key. Across poisoning attacks, adversarial queries share a rewrite-residual structure: they pair a rewrite of the target query with residual content. The rewrite maintains high similarity, while the residual elicits the malicious response. Deleting the residual makes the remaining rewrite more similar to the incoming query. We exploit this structure using Deletion Gain to search shortened variants of the cached query for similarity gains, and an Answer Check to test whether the removed text contributes to the stored answer. We prove that Deletion Gain stays positive when a deletion leaves text close enough to the rewrite, and we search for such deletions with a sliding window. Across three poisoning attack classes, our defense blocks 82.0% to 98.2% of poisoned entries at a 5% false-positive rate, with negligible serving overhead.

Comments25 pages, 5 figures, 17 tables. Code: https://github.com/shentoumengxin/deletion-gain

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑