前缀共享是一个排序问题
Prefix Sharing Is a Sorting Problem
AI总结:
本文证明LLM服务中前缀共享的片段排序问题等价于在请求上选择二叉层次结构,提出O(3^m)精确算法和1/2近似凝聚聚类,在BEIR上减少17-36%预填充。
AI中文摘要:
LLM 服务通过精确前缀匹配来复用 KV 缓存,因此当一个提示由一组可复用的片段(如检索到的段落、工具定义、少样本示例)组装而成时,这些片段被选择的顺序决定了有多少计算可以被共享。每个已部署的系统都通过单一的全局约定来固定该顺序。我们证明,仅当请求包含至多两个片段时,这种固定顺序才是最优的,而在一般情况下,其渐近表现是错误的。我们的主要结果是一个结构定理:最小前缀树代价等于 min_H ∑_x w(x) t_x(H),其中 H 是请求上的二叉层次结构,t_x(H) 是需要片段 x 的请求集合的规范分解大小。因此,选择片段顺序等价于在请求上选择一个层次结构。该恒等式产生了一个 O(3^m) 的精确算法,将两片段情形识别为最小顶点覆盖问题,并表明在留一法(leave-one-out)家族上,最优值等于二叉树的最小外部路径长度——即归并排序递归——因此全局顺序付出的代价为 Θ(n²),而真实代价为 Θ(n log n)。通过公共交集进行的凝聚聚类是实现可节省量的紧 1/2 近似。在三个 BEIR 语料库的 BM25 检索轨迹上,与生产级 RAG 排序相比,所得到的布局将预填充(prefill)减少了 17% 至 36%,并且随着检索深度的增加,差距如理论预测的那样扩大。最后,按层次结构的 DFS 顺序服务请求,使得持有单个请求上下文的缓存能够精确达到无界缓存的最优值,因此缓存容量与重排窗口互为替代。
英文摘要:
LLM serving reuses KV cache by exact prefix match, so when a prompt is assembled from a set of reusable pieces -- retrieved passages, tool definitions, few-shot exemplars -- the order chosen for those pieces determines how much computation can be shared. Every deployed system fixes that order by a single global convention. We prove this is optimal only when requests contain at most two pieces, and asymptotically wrong in general. Our main result is a structure theorem: the minimum prefix-trie cost equals min_H sum_x w(x) t_x(H) over binary hierarchies H on the requests, where t_x(H) is the canonical decomposition size of the set of requests needing chunk x. Choosing chunk orders is therefore equivalent to choosing one hierarchy over requests. The identity yields an O(3^m) exact algorithm, identifies the two-chunk case as minimum vertex cover, and shows that on the leave-one-out family the optimum is the minimum external path length of a binary tree -- the merge-sort recursion -- so a global order pays Theta(n^2) against a true cost of Theta(n log n). Agglomerative clustering by common intersection is a tight 1/2-approximation for the achievable saving. On BM25 retrieval traces over three BEIR corpora the resulting layout reduces prefill by 17-36% against production RAG ordering, and the margin widens with retrieval depth as the theory predicts. Serving requests in the hierarchy's DFS order finally lets a cache holding one request's context attain the unbounded-cache optimum exactly, so cache capacity and reorder window act as substitutes.