发表机构
University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Pallas是面向AI-RAN中LLM推理的主动KV缓存迁移框架,通过切换前在目标基站预准备推理状态,结合在线调度器优化,显著降低服务中断时间与令牌间延迟。
AI 中文摘要
AI-RAN将大语言模型(LLM)服务部署在移动用户附近,但蜂窝切换会使活跃请求与其推理状态分离:用户连接到目标基站(gNB)时,庞大且不断增长的键值(KV)缓存仍保留在源基站。在源基站保留推理可维持服务连续性,但会持续增加令牌间延迟(ITL);而在目标基站恢复状态可恢复服务局部性,但仅在切换后才需进行KV缓存传输、重新计算或两者结合,直接延长服务中断时间(SIT)。本研究提出Pallas,一种主动KV缓存迁移框架,在切换前于预测的目标基站准备推理状态,与源侧推理和令牌传输并行进行。在准备触发时,Pallas将令牌序列划分为稳定的历史前缀和不断演化的后缀。目标基站通过本地预填充重构前缀,源基站则流式传输为后缀生成的KV块。切换时,目标基站将两部分组装为最新的KV缓存并在本地恢复解码,仅留下未完成的准备工作导致SIT。在线调度器基于移动性预测和运行时遥测选择预取窗口,该窗口决定切换前准备开始的时间。在三个LLM和100-500 Mbps的gNB间链路上,我们基于vLLM的原型相比目标侧恢复方法将平均SIT降低2.28至89.68倍,相比源侧转发将平均ITL降低16.0%至50.0%。
英文摘要
AI-RAN brings large language model (LLM) serving close to mobile users, but cellular handover can separate an active request from its inference state: the user attaches to a target base station (gNB) while the large and growing key-value (KV) cache remains at the source. Retaining inference at the source preserves service continuity but persistently increases inter-token latency (ITL), whereas recovering the state at the target restores serving locality but requires KV-cache transfer, recomputation, or a combination of both only after handover, directly prolonging service interruption time (SIT). This work presents Pallas, a \textit{proactive} KV-cache migration framework that prepares the inference state at the predicted target before handover, in parallel with ongoing source-side inference and token delivery. At the preparation trigger, Pallas partitions the token sequence into a stable historical prefix and an evolving suffix. The target reconstructs the prefix through local prefill, while the source streams the KV blocks generated for the suffix. At handover, the target assembles both portions into an up-to-date KV cache and resumes decoding locally, leaving only unfinished preparation to contribute to SIT. An online scheduler selects the \textit{prefetching window}, which determines how early preparation begins before handover, based on mobility predictions and runtime telemetry. Across three LLMs and $100$--$500~\mathrm{Mbps}$ inter-gNB links, our vLLM-based prototype reduces average SIT by factors of $2.28$--$89.68$ over target-side recovery approaches and lowers average ITL by $16.0\%$--$50.0\%$ compared with source-side forwarding.