PLoRA:一种用于高性价比多LoRA服务的NDP增强型池化内存系统
PLoRA: An NDP-Enhanced Pooled-Memory System for Cost-Efficient Multi-LoRA Serving
AI总结:
PLoRA是一种NDP增强型池化内存系统,通过优化适配器与KV缓存的存储和访问方式,在H100上服务1000个适配器时,解码延迟较S-LoRA平均低6.6倍,实现了高性价比的多LoRA服务。
AI中文摘要:
多LoRA服务是指一个基础模型衍生出数千种专用变体,每个变体对应一个用户、任务或智能体的适配器,部署中可容纳1000个以上适配器。对这些适配器进行服务部署存在诸多挑战:其工作负载与GPU的资源特性相反,GPU具备数十万亿次浮点运算性能,但多LoRA服务需要数TB级内存;此外,现有所有已发布的系统均通过PCIe将适配器从CPU DRAM加载,每次访问都会产生内核暂停和主机运行复制的开销,且容量受限于主板的DIMM插槽。与此同时,内存语义 fabrics(如CXL和NVLink)正逐步向池化内存演进,加速器可通过自身加载和存储指令访问该池化内存,而近数据处理(NDP)技术可将计算单元部署在池化数据旁。如何在这类硬件上服务多LoRA工作负载仍是未被探索的问题。本文提出PLoRA,一种用于高性价比多LoRA服务的NDP增强型池化内存系统。PLoRA将适配器和KV缓存存储在池中,仅通过GPU驱动的读取-计算接口在链路上返回简化结果。在该架构之上,GPU内存管理系统会为每个适配器从4种LoRA和2种注意力执行策略中选择,并基于链路参数化成本模型将性能最关键的字节缓存至GPU内存中。在一台H100上服务1000个适配器时,PLoRA在所有测试的模型和工作负载上均实现了最低的解码延迟,平均比真实机器的S-LoRA低6.6倍,且设备面积仅增加不到3.4%。链路本身的影响被消除:在短上下文下吞吐量在32GB/s时达到饱和,仅为CXL 3.1的四分之一;且该结论在扩展场景下依然成立:当适配器流量与张量并行分片后,每GPU需求从70亿参数级降至建模的1.2万亿参数级部署。该设计可在从CXL级到NVLink级的fabrics上无修改运行,剩余带宽可用于增加池化容量而非提升速度。
英文摘要:
Multi-LoRA serving is how one base model becomes thousands of specialized variants, one adapter per user, task, or agent, and the deployments can hold 1000-plus adapters. Serving them is hard because the workload inverts what GPUs provide: terabytes of memory against only tens of TFLOPS, and because every published system stages its adapters from CPU DRAM over PCIe, where each access pays a kernel stop and a host-run copy and capacity ends at the motherboard's DIMM slots. Meanwhile, memory-semantic fabrics such as CXL and NVLink are converging on pooled memory that an accelerator addresses with its own loads and stores, and near-data processing (NDP) can place compute beside the pooled data. How to serve multi-LoRA workloads on such hardware remains unexplored. This paper introduces PLoRA, an NDP-enhanced pooled-memory system for cost-efficient multi-LoRA serving. PLoRA keeps adapters and KV cache in the pool and returns only reduced results over the link, through a read-compute interface the GPU drives with its own loads and stores. Above this architecture, a GPU memory management system picks among four LoRA and two attention execution strategies for each adapter and caches the most performance-critical bytes in GPU memory, guided by a link-parameterized cost model. On one H100 serving 1000 adapters, PLoRA attains the lowest decode latency on every model and workload we measure, averaging 6.6x below a real-machine S-LoRA at under 3.4% added device area. The link itself stops mattering: throughput saturates at 32 GB/s on short contexts, a quarter of CXL 3.1, and the verdict survives scale: per-GPU demand falls from 7B to a modeled 1.2T deployment once adapter traffic shards with the tensor parallelism. The design runs unchanged from CXL-class to NVLink-class fabrics, and surplus bandwidth buys pooled capacity rather than speed.