用于消费级设备上混合CPU - GPU大语言模型推理的自动张量调度
Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices
浏览论文内容
中文总结 AI 辅助
研究在消费级设备运行大语言模型时因模型权重超GPU内存需卸载推理的问题,提出ATSInfer系统,通过结合静态张量放置与负载感知动态传输及异步协调,在代表性平台评估,相比现有系统显著提升吞吐量、利用率等,改善本地大语言模型部署体验。
中文摘要 AI 辅助
在笔记本电脑和台式机等消费级设备上运行大语言模型具有挑战性,因为模型权重常超GPU内存容量,需借助CPU内存进行卸载推理。现有卸载系统依赖粗略调度,忽略张量异质性且难适应硬件负载变化。本文提出ATSInfer,一种在张量粒度上执行卸载的混合CPU - GPU推理系统。它结合静态张量放置与负载感知动态传输,引入异步CPU - GPU协调,跨异构后端高效调度硬件存储、数据移动和计算。在代表性消费平台上用密集和MoE模型评估,相比现有系统,ATSInfer提高预填充吞吐量达1.94倍,解码吞吐量达3.29倍,还提升了GPU利用率并更有效利用PCIe带宽,显著改善了个人消费设备上本地大语言模型部署的用户体验。
英文摘要
Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory capacity, making offloading inference necessary to extend effective model capacity with CPU memory. Existing offloading systems, however, typically rely on coarse layer-level or expert-level scheduling, which overlooks substantial heterogeneity among tensors within the same layer and adapts poorly to changing hardware load conditions on such devices. This paper presents ATSInfer, a hybrid CPU-GPU inference system for consumer devices that performs offloading at tensor granularity. ATSInfer combines static tensor placement with load-aware dynamic transfer, and introduces asynchronous CPU-GPU coordination to efficiently schedule hardware storage, data movement, and computation across heterogeneous backends. We implement ATSInfer and evaluate it on representative consumer platforms using both dense and MoE models. Compared with existing systems, ATSInfer improves prefill throughput by up to 1.94$\times$ and decode throughput by up to 3.29$\times$, while also increasing GPU utilization and making more effective use of PCIe bandwidth. These results show that ATSInfer can substantially improve the user experience of local LLM deployment on personal consumer devices.
发表机构
- School of Computer Science, Nanjing University(南京大学计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。