arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型(LLM)缓存应采用哪种驱逐策略?跨工作负载、容量和编码器的系统研究

Which Eviction Policy Should an LLM Cache Use? A Systematic Study Across Workloads, Capacities, and Encoders

Yash Kulkarni, Shubham Harkare, Arvind Yogesh Suresh Babu

arXiv 2608.20280首次发表:更新:

发表机构

University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究使用CLEVER工具在多场景下对比多种LLM缓存驱逐策略,发现LFU性能最优,且当前操作点的答案可替换性低、阈值难跨模型迁移,建议部署时先验证答案有效性再测试次级策略差异。

AI 中文摘要

语义缓存会在传入查询的嵌入与已缓存查询相近时复用大语言模型(LLM)的响应,但各类已提出的驱逐策略很少在同一协议下进行对比。本研究使用CLEVER工具,在三个有序且去重的查询语料库、三种缓存容量及两种编码器(包括MiniLM)的场景下,对FIFO、LRU、LFU、ARC、GDSF、SISO的单遍流式适配版本,以及一种基于语义冗余的策略进行评估。在全部18种设置中,无任何被评估策略的性能超过LFU达0.041个百分点以上。替换策略并非无关紧要:在缓存容量紧张时,FIFO和流式SISO的表现分别比LFU低达8.67和8.55个百分点。本研究通过条件打包结果解释了缺失的性能提升:在精确查找和未命中时插入的规则下,新插入的条目不可能在命中半径内存在已驻留的邻居,因此感知几何的驱逐规则几乎无法获得新的冗余信号。一项独立审计还揭示了评估操作点存在更大问题:在MiniLM的中位数最近邻阈值下,仅2.1%-3.9%的采样LMSYS和QQP命中被判定为可替换答案,将51%-60%的原始命中率降至质量调整后的1.1%-2.2%。跨编码器研究进一步表明,阈值无法在不同嵌入模型间迁移。本研究得出结论:在该协议下,LFU是最强的简单默认策略;部署决策应先确定答案有效性,再通过精确搜索测试次级策略的差异。

英文摘要

Semantic caches reuse an LLM response when the incoming query embedding lies near a cached query, but proposed eviction policies have rarely been compared under one protocol. Using CLEVER, we evaluate FIFO, LRU, LFU, ARC, GDSF, a single-pass streaming adaptation of SISO, and a semantic-redundancy policy across three ordered, deduplicated query corpora, three cache capacities, and two encoders. No evaluated policy improves on LFU by more than 0.041 percentage points in any of the eighteen settings. Replacement is not irrelevant: FIFO and streaming SISO trail LFU by as much as 8.67 and 8.55 points, respectively, at tight capacity. We explain the missing upside with a conditional packing result. Under exact lookup and insert-on-miss, a newly inserted entry cannot have a resident neighbor within the hit radius, so a geometry-aware eviction rule receives little new redundancy signal. A separate audit exposes a larger problem with the evaluated operating point. At MiniLM's median nearest-neighbor threshold, only 2.1-3.9% of sampled LMSYS and QQP hits are judged answer-substitutable, reducing raw hit rates of 51-60% to quality-adjusted rates of 1.1-2.2%. The cross-encoder study further shows that thresholds do not transfer between embedding models. LFU is the strongest simple default in this protocol; deployment decisions should first establish answer validity and then test sub-point policy differences with exact search.

Comments11 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑