ImpactHO:面向多用户边缘大模型切换的重要性感知键值缓存迁移
ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover
- Pohang University of Science and Technology (POSTECH)(浦项科技大学(POSTECH))
- Ajou University(明知大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对多用户边缘大模型切换时回程链路饱和导致缓存无法完整迁移的问题,提出重要性感知KV缓存迁移方案,通过闭式加权注水算法实现高效分配,在500ms窗口内达到接近全缓存的准确率。
AI中文摘要:
边缘大模型(Edge LLMs)在用户在边缘节点间切换时必须保持推理连续性,这需要将键值(KV)缓存迁移到目标节点。但同时切换会使回程链路饱和,无法在移动性强制的迁移窗口内完成全部缓存传输。我们未像所有缓存条目价值相等那样分配带宽,而是按重要性排序每个用户的KV缓存,仅传输其最具信息性的部分,将令牌级稀疏性转化为通信节省。我们将迁移问题建模为最大化用户平均准确率的多用户回程链路分配问题,每个用户的部分缓存准确率作为其效用:这是一个sigmoid函数,在RULER基准测试中,对不同模型和上下文长度的拟合R²均大于0.99。由于重要性排序将高价值条目前置,准确率曲线的凹区域几乎覆盖整个缓存。我们提出的分配器将服务用户保持在该区域内,使每个时隙分配问题变为凸问题。最优解通过闭式加权注水算法得到,该算法推广了信息论中的注水算法并支持在线调度。在500ms的迁移窗口中,所提分配器达到超过93.7%的平均准确率,与全缓存上限相差在0.5个百分点内,达到全知上限的98.2%-99.5%。
英文摘要:
Edge LLMs must preserve inference continuity when a user hands over between edge nodes, requiring key-value (KV) cache transfer to the target node. However, simultaneous handovers saturate the backhaul, preventing full cache delivery within the mobility-imposed transfer window. Rather than allocating bandwidth as if all cache entries were equally valuable, we order each user's KV cache by importance and transmit only its most informative fraction, turning token-level sparsity into communication savings. We cast the transfer as a multi-user backhaul allocation problem that maximizes average accuracy across users. Each user's partial-cache accuracy serves as its utility: a sigmoid that fits measurements on the RULER benchmark with $R^2>0.99$ across models and context lengths. Because importance ordering front-loads the high-value entries, the concave region of the accuracy curve spans nearly the entire cache. Our proposed allocator keeps served users within this region, making each per-slot allocation problem convex. The optimum is derived via a closed-form weighted water-filling solution that generalizes information-theoretic water-filling and enables online scheduling. The proposed allocator attains over 93.7% average accuracy in a 500ms transfer window, within 0.5pp of the full-cache ceiling, and reaches 98.2-99.5% of a clairvoyant upper bound.