发表机构
Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Beaver通过动态调整SM分配和重写ML内核,在共享GPU上保护vRAN延迟关键型工作负载的截止时间,同时保留74%的LLM服务吞吐量。
AI 中文摘要
在AI-RAN的行业共同愿景中,AI与RAN旨在将虚拟化无线接入网(vRAN)工作负载和AI服务共同部署在共享GPU上。这种共享本质上是不对称的:vRAN工作负载是延迟关键型的,而机器学习(ML)工作负载则是面向吞吐量的、尽力而为的共租户。我们提出了Beaver,一个GPU共享系统,它联合管理计算和内存资源,同时保护vRAN的严格处理截止时间。Beaver根据每个时隙的调度工作负载来调整vRAN的流式多处理器(SM)分配,在时隙粒度上重新划分SM分配,并重写编译后的ML内核,以在vRAN的内存关键阶段腾出HBM带宽。我们实现了Beaver,并使用NVIDIA Aerial、异构多小区工作负载、真实蜂窝网络轨迹、生产推理内核以及全栈LLM服务对其进行了评估。在H200 GPU上,Beaver将vRAN的p99.9延迟保持在1.5ms上行链路截止时间之内,同时保留了Llama-3.3-70B服务吞吐量的74%。在重放的蜂窝网络轨迹下,它也没有出现可观察到的截止时间错过,保护了375us的下行链路截止时间,并且可推广到其他GPU,包括A100、GB10和GH200。
英文摘要
Within the shared industry vision of AI-RAN, AI-and-RAN seeks to co-locate virtualized radio access network (vRAN) workloads and AI services on shared GPUs. This sharing is inherently asymmetric: vRAN workload is latency-critical, whereas the machine learning (ML) workload is a throughput-oriented, best-effort co-tenant. We present Beaver, a GPU sharing system that jointly manages compute and memory resources while protecting the vRAN's strict processing deadline. Beaver sizes the vRAN's streaming multiprocessor (SM) allocation from each slot's scheduled workload, repartitions SM allocations at slot granularity, and rewrites compiled ML kernels to yield HBM bandwidth during the vRAN's memory-critical phases. We implement Beaver and evaluate it using NVIDIA Aerial with heterogeneous multi-cell workloads, real-world cellular traces, production inference kernels, and full-stack LLM serving. On an H200 GPU, Beaver keeps the vRAN's p99.9 latency within its 1.5ms uplink deadline while retaining 74% of Llama-3.3-70B serving throughput. It also incurs no observed deadline misses under replayed cellular traces, protects a 375us downlink deadline, and generalizes to other GPUs including A100, GB10 and GH200.
Comments16 pages, 16 figures