LaCache:面向大语言模型服务的鲁棒语义缓存
LaCache: Robust Semantic Caching for LLM Serving
浏览论文内容
中文总结 AI 辅助
LaCache是一种新型语义缓存方案,通过检查查询及其前k个解码标记的缓存命中,可抵御缓存碰撞攻击,提升响应相关性,经多种大语言模型及基准验证了其安全性与效率优势。
中文摘要 AI 辅助
语义缓存通过嵌入向量复用语义相似请求的响应,在大语言模型(LLM)服务中应用日益广泛,可实现更快响应并降低成本。然而,现有方案存在根本性缺陷,易受缓存碰撞攻击:攻击者注入精心构造的查询污染缓存,会损坏后续合法请求的响应。本文提出LaCache,一种新型语义缓存方案,通过概念简单且原则性强的重新设计解决该漏洞。核心见解是,攻击者虽能完全控制对抗性查询,但对其响应的控制能力有限,响应需同时满足多个语义约束。LaCache不仅检查查询的缓存命中,还额外检查其前k个(推测)解码标记的缓存命中。该设计带来两个具体好处:一是提供形式化保证的缓存碰撞攻击抵御能力,证明无法构造能同时引发恶意响应并与良性查询碰撞的对抗性查询;二是丰富的索引为缓存检索提供额外语义上下文,提升响应相关性。对多种大语言模型及基准的实证评估验证了LaCache的安全保证与效率提升,为鲁棒语义缓存指明了有前景的方向。
英文摘要
Semantic caching, which reuses responses to semantically similar requests via their embeddings, has seen growing adoption in LLM serving, offering faster responses and reduced costs. Yet existing schemes are fundamentally vulnerable to cache-collision attacks, wherein an adversary pollutes the cache by injecting crafted queries, corrupting responses to subsequent legitimate requests. We present LaCache, a novel semantic caching scheme that addresses this vulnerability through a conceptually simple yet principled redesign. The key insight is that while the adversary has full control over the adversarial query, it has far less control over its response, which must simultaneously satisfy multiple semantic constraints. Rather than checking only the cache hit of a query, LaCache additionally checks the cache hit of its first k (speculatively) decoded tokens. This design yields two concrete benefits. First, it provides formally guaranteed resilience against cache-collision attacks: we prove that it is impossible to craft adversarial queries that simultaneously elicit malicious responses and collide with benign queries. Second, the enriched index supplies additional semantic context for cache retrieval, improving response relevance. Empirical evaluation across diverse LLMs and benchmarks validates both LaCache's security guarantees and efficiency gains, pointing to a promising direction for robust semantic caching.