面向无线边缘大语言模型推理的空间前缀缓存:一种随机几何与排队框架
Spatial Prefix Caching for Wireless Edge LLM Inference: A Stochastic-Geometry and Queueing Framework
浏览论文内容
中文总结 AI 辅助
本文构建随机几何与排队框架,针对无线边缘LLM推理的空间通信-缓存-计算权衡优化,揭示延迟最优节点非最近节点、TTFT随缓存前缀深度非单调变化的规律。
中文摘要 AI 辅助
前缀缓存可复用共享提示前缀的键值(KV)状态,能大幅缩短大语言模型(LLM)推理的首token生成时间(TTFT)。然而在无线边缘网络中,前缀状态分布在地理上分离的GPU节点上:邻近节点的无线路径短,但复用率低;更远的节点可能缓存更长的匹配前缀,但会产生额外的通信和排队延迟。此外,持久前缀和活跃请求的KV状态会竞争相同的GPU内存,因此激进的缓存会降低推理并发度,形成排队热点。本文针对这种空间通信-缓存-计算的权衡,构建了随机几何与排队框架:将提示工作负载表示为前缀森林,定义可包含多个可复用前缀的祖先封闭缓存配置文件;边缘GPU节点形成泊松点过程,由缓存配置文件独立标记,产生可解析处理的空间层级。推导了负载感知关联策略下的配置文件关联概率、条件服务距离分布、token级计算卸载率和TTFT覆盖概率;用不动点公式刻画空间关联与多服务器GPU队列的耦合关系,通过外部优化在静态内存和稳定性约束下选择缓存配置文件分布。分析结果与蒙特卡洛结果高度吻合,显示延迟最优节点未必是最近节点,且受GPU内存耦合影响,TTFT随缓存前缀深度呈非单调变化。
英文摘要
Prefix caching reuses the key--value (KV) states of shared prompt prefixes and can substantially reduce the time to first token (TTFT) of large language model (LLM) inference. In a wireless edge network, however, prefix states are distributed across geographically separated GPU nodes. A nearby node offers a short radio path but may provide little reuse, whereas a more distant node may cache a longer matching prefix but incur additional communication and queueing delay. Moreover, persistent prefixes and active-request KV states compete for the same GPU memory, so aggressive caching can reduce inference concurrency and create queueing hotspots. This paper develops a stochastic-geometry and queueing framework for this spatial communication--caching--computation tradeoff. We represent the prompt workload by a prefix forest and define an ancestor-closed cache profile that may contain multiple reusable prefixes. Edge GPU nodes form a Poisson point process and are independently marked by cache profile, yielding analytically tractable spatial tiers. We derive the profile-association probability, conditional serving-distance distribution, token-level computation-offloading ratio, and TTFT coverage probability under a load-aware association policy. A fixed-point formulation captures the coupling between spatial association and multi-server GPU queues, while an outer optimization selects the cache-profile distribution subject to static-memory and stability constraints. Analytical and Monte Carlo results agree closely. The results show that the latency-optimal node need not be the nearest node, that TTFT can be non-monotonic in cached-prefix depth because of GPU-memory coupling.