发表机构
UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EMA通过跨GPU弹性内存共享与预取机制,实现性能透明,提升用户吞吐量最高52%,并保持接近2倍容量系统的性能。
AI 中文摘要
多GPU服务器已成为现代数据中心的标准构建模块,通过高带宽互连提供聚合容量。与此同时,LLM推理等工作负载表现出高度动态的内存需求,这可能导致一个GPU耗尽本地内存,而其他GPU仍未被充分利用。这种不匹配促使了跨GPU弹性资源共享模型的产生。我们提出了EMA,一个内存共享系统,允许服务器内的GPU相互借用和回收内存,形成一个弹性的容量池。EMA确保借用方和出借方都能获得性能透明性。对于借用方,预取隐藏了远程访问成本,使得应用程序在性能上无法区分远程内存和本地内存。对于出借方,借出的资源可按需回收,保证性能永远不会低于静态分区时的水平。虽然我们的设计聚焦于内存,但同样的原理自然扩展到其他GPU资源。我们的评估显示,EMA将单个用户吞吐量提升了高达52%,达到了配备2倍容量系统吞吐量的96%,并保持了与静态本地基线相似的延迟。
英文摘要
Multi-GPU servers have become the standard building block of modern data centers, providing aggregated capacity through high-bandwidth interconnects. At the same time, workloads such as LLM inference exhibit highly dynamic memory demands, which can cause one GPU to exhaust its local memory while others remain underutilized. This mismatch motivates a model of elastic resource sharing across GPUs. We present EMA, a memory sharing system that allows GPUs within a server to borrow and reclaim memory from each other, forming an elastic pool of capacity. EMA ensures performance transparency for both borrowers and lenders. For borrowers, prefetching hides remote access costs so that applications experience remote and local memory as indistinguishable in performance. For lenders, borrowed resources remain reclaimable on demand, guaranteeing that performance never falls below that of static partitioning. While our design focuses on memory, the same principle naturally extends to other GPU resources. Our evaluation shows that EMA improves individual user throughput by up to 52%, achieves 96% of the throughput of a system provisioned with 2X capacity, and maintains latency similar to the static local baseline.