arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

mzCache:多任务场景下的设备端大语言模型内存管理

mzCache: On-Device LLM Memory Management under Multitasking

Hongseung Yu, Minsung Kim, Jongseok Park, Kyunghan Lee

arXiv 2609.01338首次发表:更新:

发表机构

Seoul National University; UC Berkeley(首尔大学; 加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

mzCache是面向多任务场景的设备端LLM内存管理系统,通过弹性回收、面向恢复的内存管理及混合策略,实现低延迟推理,首token生成时间较存储卸载方案缩短2.1-5.5倍。

AI 中文摘要

设备端移动大语言模型(LLM)推理正受到广泛关注。然而,移动设备运行在高度动态的多任务环境中,用户会频繁在应用间切换,这会产生内存压力,迫使操作系统回收LLM内存(模型权重与KV缓存)。当新的推理请求到达时,推理系统必须通过慢速存储读取恢复被回收的内存,或重新计算整个KV缓存,严重降低响应速度。为解决该问题,本文提出mzCache,一款面向多任务环境的设备端LLM推理系统,具备专用内存管理能力。在不可预测的内存压力下,mzCache弹性回收LLM内存,并利用移动SoC的统一内存,通过CPU侧并发恢复实现GPU的零等待推理。该系统通过面向恢复的内存管理实现上述功能:将LLM内存划分为细粒度共享缓冲区,支持部分回收与恢复,同时支持跨处理器并发访问;混合交换与反向回收策略确保从任意回收状态实现低延迟恢复。mzCache在[该链接]实现并部署为Android应用,与基于存储的部分卸载方案相比,其首token生成时间缩短2.1至5.5倍,在真实多任务场景中验证了有效性。

英文摘要

On-device mobile Large Language Model (LLM) inference is gaining significant attention. However, mobile devices operate in highly dynamic multitasking environments where users frequently switch between applications. This creates memory pressure, forcing LLM memory (model weights and KV cache) to be evicted by the operating system. When a new inference request arrives, the inference system must restore the evicted memory through slow storage reads or recompute the entire KV cache, severely degrading responsiveness. To address this, we present mzCache, an on-device LLM inference system with specialized memory management for multitasking environments. Under unpredictable memory pressure, mzCache elastically evicts LLM memory and leverages the unified memory of mobile SoCs to enable zero-wait inference on the GPU with concurrent CPU-side restoration. mzCache realizes this through restoration-oriented memory management: LLM memory is partitioned into fine-grained shared buffers to enable partial eviction and restoration with concurrent cross-processor access, while hybrid swap and backward-out eviction policies ensure low-latency restoration from any eviction state. Implemented on llama.cpp and deployed as an Android application, mzCache achieves 2.1-5.5$\times$ reduction in Time-to-First-Token compared to storage-backed partial offload and demonstrates its effectiveness in real multitasking scenarios.

CommentsMobiCom 2026

DOI:10.1145/3795866.3844495

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑