arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

动态流,静态图:面向移动NPU的高效LLM服务的KV缓存复用

Dynamic Flow, Static Graph: KV Cache Reuse for Efficient LLM Serving on Mobile NPUs

Zhengxiang Huang, Shengheng Chen, Chaoyue Niu, Yujie Sun, Zhaode Wang, Zeyu Zhao, Chengfei Lv, Fan Wu, Guihai Chen

arXiv 2609.34727首次发表:更新:

发表机构

Shanghai Jiao Tong University; Alibaba Group(上海交通大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对移动NPU静态图与带宽限制,提出计算-存储协同的KV缓存复用设计,通过图内重计算、图间调度及分层管理,将TTFT降低40-60%。

AI 中文摘要

设备端大语言模型(LLM)服务是本地优先个人智能的基石,为用户提供数据主权、强大的隐私保证,并摆脱云API延迟和成本。尽管KV缓存被广泛用于减少长上下文推理中的延迟,但现有设计主要针对具有动态执行环境和充足内存带宽的云GPU进行优化。这些架构假设在移动NPU上不成立,因为移动NPU的计算图必须静态编译,且内存容量和I/O带宽都受到严重限制。在这项工作中,我们提出了一种面向移动端前缀和非前缀KV复用的计算-存储协同设计。我们首先提出一种图内机制,将选择性KV重计算映射到静态NPU图上,协调算法动态性与NPU静态性。我们进一步开发了一种图间调度器,通过动态规划优化块合并并最小化填充。为了解决移动带宽限制,我们引入了一种分层KV管理器,具有树-哈希-语义混合结构,以及成本感知的预取和驱逐策略。我们还构建了一个二维流水线,将KV加载、旋转和存储与NPU执行重叠,隐藏数据移动延迟。在代表性设备端工作负载和LLM上的实验表明,与无复用和仅前缀缓存相比,我们的设计将首令牌时间(TTFT)降低了40-60%。

英文摘要

On-device large language model (LLM) serving is a cornerstone of local-first personal intelligence, offering users data sovereignty, strong privacy guarantees, and freedom from cloud API latency and cost. Although KV caching is widely used to reduce latency in long-context inference, existing designs were primarily optimized for cloud GPUs with dynamic execution environments and abundant memory bandwidth. These architectural assumptions do not hold on mobile NPUs, where computation graphs must be statically compiled and both memory capacity and I/O bandwidth are severely constrained. In this work, we present a compute-storage co-design for mobile-centric prefix and non-prefix KV reuse. We first propose an intra-graph mechanism that maps selective KV recomputation onto static NPU graphs, reconciling algorithmic dynamicity with NPU staticity. We further develop an inter-graph scheduler to optimize chunk merging and minimize padding with dynamic programming. To address mobile bandwidth limitations, we introduce a hierarchical KV manager featuring a tree-hash-semantic hybrid structure, along with cost-aware prefetching and eviction policies. We also build a two-dimensional pipeline that overlaps KV loading, rerotation, and storage with NPU execution, hiding data-movement latency. Experiments across representative on-device workloads and LLMs show that our design reduces time-to-first-token (TTFT) by $40-60\%$ compared with no reuse and prefix-only caching.

CommentsAccepted at EuroSys 2027 (Spring Cycle)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑