arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24204cs.NI

WiCi:无线GPU计算基础设施

WiCi: Wireless GPU Computing Infrastructure

Yibin Shen, Wei Li, Kaiqiang Xu, Zili Meng

中文总结 AI 辅助

针对移动设备本地推理性能不足、云端推理成本高的问题,本文提出WiCi无线GPU计算基础设施,可让移动设备通过WiFi接入服务器级GPU,大幅提升推理性能并支持更大模型,性能接近服务器级GPU原生表现。

中文摘要 AI 辅助

大语言模型(LLM)推理应用正获得显著发展,推理需求呈指数级增长,推理任务的GPU使用量已逐渐超过训练任务。由于移动性带来的性能损失,边缘侧推理无法提供令人满意的表现,因此目前大多数推理服务提供商依赖基于云端的推理,这给企业带来了巨大且难以持续的成本,在智能体(agentic)范式下该成本甚至还在上升。因此,本文的目标是在移动设备上实现服务器级GPU的强大计算能力,我们提出了无线GPU计算基础设施(WiCi)。通过WiCi,移动设备可无线接入服务器级GPU,在移动客户端运行推理任务,同时将与GPU相关的计算卸载到附近的GPU上,且该过程通过WiFi实现。WiCi引入了一系列设计,以确保该基础设施可适配不同应用、兼容不同移动设备,且性能可与在物理GPU上运行的表现相当。我们从移动设备对WiCi进行测试,结果显示:对于同一模型,与移动设备本地推理相比,WiCi可将首token生成时间最多降低90%,token生成速率提升约39倍,还支持更大规模的模型;在不同应用场景下,WiCi的性能也达到了服务器级GPU原生性能的近80%。

英文摘要

LLM inference applications are gaining significant traction. The demand for inference is growing exponentially, and the GPU usage of inference is increasingly surpassing that of training. Due to the mobility penalty, edge-side inference fails to deliver satisfactory performance. Consequently, most inference service providers currently rely on cloud-based inference, which incurs substantial, not sustainable costs for enterprises, and is even increasing in the agentic paradigm. Therefore, our goal is to enable powerful computing capabilities as server-grade GPUs on mobile devices. We propose Wireless GPU Computing Infrastructure (WiCi) in this paper. Through WiCi, mobile devices can wirelessly access server-grade GPUs, running inference tasks on mobile clients but offloading GPU-related computations to a nearby GPU via WiFi. WiCi introduces a series of designs to make sure the infrastructure is scalable with different applications, compatible with different mobile devices, and has comparable performance to running on a physical GPU. We test WiCi from mobile devices and find that WiCi can reduce time to first token by up to 90%, improve the token rate by approximately 39x compared to local inference on mobile devices for the same model, and support much larger models. WiCi also achieves up to nearly 80% of the native performance of the server-grade GPU across different applications.

↑