Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
在分布式大语言模型推理系统中服务异构LoRA适配器
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Microsoft(微软) ; National Technical University of Athens(雅典国家技术大学)
AI总结 LoRAServe通过动态适配器放置和路由框架,有效解决异构LoRA适配器在分布式大语言模型推理中的性能偏斜问题,提升吞吐量并降低延迟。