arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AIR-LLM:通过射频计算在无线电上广播AI权重,实现无内存边缘LLM推理

AIR-LLM: Broadcasting AI Weights over Radio for Memory-Free Edge LLM Inference via RF Computing

Zhihui Gao, Tingjun Chen, Dirk Englund

arXiv 2610.00465首次发表:更新:

发表机构

Massachusetts Institute of Technology; Duke University(麻省理工学院; 杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出AIR-LLM架构,通过无线广播LLM权重并在射频域直接计算,实现边缘设备无内存推理,显著降低能耗和空中时间。

AI 中文摘要

下一代大语言模型(LLM)正从云端扩展到无处不在的边缘设备。然而,边缘设备通常要么缺乏存储日益庞大的LLM权重的内存,要么即使内存充足,也会在加载权重上花费难以承受的能量。这引出了我们的问题:边缘设备能否在不存储或加载权重的情况下运行LLM,而是通过空中接收权重并即时消耗它们?受无线广播的启发,我们提出了AIR-LLM,一种面向边缘设备的LLM推理架构,它由以下部分组成:(i)一个中央无线电(例如5G基站),将LLM权重广播到空中;(ii)边缘用户接收权重,并直接在射频(RF)域中使用RF混频器完成LLM推理的通用矩阵-向量乘法(GEMV)。为了进一步缩短空中时间,AIR-LLM利用MIMO空间复用,并提出了一种在边缘侧的能量高效的预编码-后编码器对,以校准其自身的无线信道。由于中央无线电保持用户无关性,AIR-LLM具有用户可扩展性,一次广播即可服务覆盖范围内的无限用户。我们在两个真实城市场景的NVIDIA Sionna光线追踪信道和真实RF混频器的性能分析上实现了AIR-LLM。在WikiText-2困惑度下降4.0%(相对于LLaMA-3.1-8B)的情况下,AIR-LLM相比FP16和仅权重量化基线分别节省了157.7倍和40.4倍的能量;在20个用户的情况下,其空中时间分别缩短了104.1倍和26.0倍。

英文摘要

Next-generation large language models (LLMs) are expanding from the cloud to ubiquitous edge devices. However, edge devices typically either lack the memory to store increasingly large LLM weights or, even with enough memory, spend unaffordable energy on loading the weights. This raises our question: can an edge device run an LLM without storing or loading its weights, but receive them over the air and consume them on the fly? Inspired by wireless broadcasting, we present AIR-LLM, an LLM inference architecture for edge devices, which is composed of: (i) a central radio (e.g., 5G base stations) that broadcasts the LLM weights into the air, and (ii) the edge user that receives the weights and completes the general matrix-vector multiplication (GEMV) of LLM inference directly in the radio frequency (RF) domain using RF mixers. To further shorten the airtime, AIR-LLM exploits MIMO spatial multiplexing and proposes an energy-efficient precoder-postcoder pair on the edge to calibrate its own wireless channel. Since the central radio stays user-unaware, AIR-LLM is user-scalable so that one broadcast serves unlimited users within its coverage. We implement AIR-LLM on the NVIDIA Sionna ray-traced channels of two real-world urban scenes and the profiling of a real RF mixer. With a WikiText-2 perplexity degradation of 4.0% on LLaMA-3.1-8B, AIR-LLM saves the energy by 157.7x/40.4x against the FP16 and weight-only quantization baselines; with 20 users, its airtime is 104.1x/26.0x shorter, respectively.

Comments14 pages, 12 figures, 6 tables. Appendix: 12 pages, 7 figures, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑