发表机构
Shanghai Jiao Tong University; Lenovo Information Products (Shenzhen) Ltd; Lenovo (Beijing) Ltd(上海交通大学; 联想信息产品(深圳)有限公司; 联想(北京)有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在实现局域网内跨设备GPU共享以促进AI推理,提出Gleam框架,通过自动模型权重缓存、异步执行、设计运行时任务调度器及确保上下文一致性等方法,提升API远程调用效率和系统吞吐量,性能优于现有基线。
AI 中文摘要
本文旨在实现局域网内跨设备的高效计算和通信GPU共享,促进异构个人设备上的普遍AI推理。通过CUDA API远程调用实现分布式任务卸载,但网络限制成为主要瓶颈。为此提出Gleam框架,有三个关键贡献:通过自动模型权重缓存减少带宽开销,异步执行减轻频繁API调用的延迟;设计运行时任务调度器动态确定API远程调用对;引入专用机制确保分布式执行中CUDA上下文一致性。实验表明Gleam性能优于现有基线,API远程调用效率提高1.4 - 24.2倍,系统吞吐量提高1.79倍。
英文摘要
This paper aims to enable computation- and communication-efficient GPU sharing across devices within local area networks (LANs), facilitating ubiquitous AI inference on heterogeneous personal devices. We achieve distributed task offloading via CUDA API remoting. However, beyond raw computation, network constraints emerge as the primary bottleneck: limited bandwidth, high-frequency API invocations, and cross-task contention significantly hinder performance. To address these challenges, we propose Gleam, a novel and network-efficient framework for task-generic GPU sharing across local-area CUDA devices, with three key contributions. First, we reduce bandwidth overhead in CUDA API remoting through automatic model weight caching, and mitigate accumulated latency from frequent API calls by asynchronous execution. Second, we design a runtime task scheduler that dynamically determines API remoting pairs between LAN clients and servers, explicitly accounting for both network conditions and GPU resource contention under parallel workloads. Finally, we introduce dedicated mechanisms to ensure CUDA context consistency across distributed executions. Extensive experiments on heterogeneous NVIDIA GPUs and diverse AI workloads show Gleam consistently outperforms state-of-the-art baselines, achieving 1.4-24.2 times improvements in API remoting efficiency and up to 1.79 times higher system throughput.
Comments20 pages, 28 figures