野外环境下的大语言模型服务:框架、方法与系统设计的实证研究
LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs
- Polytechnique Montréal(蒙特利尔理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过实证分析开源软件系统中5种LLM服务框架的应用情况,明确了框架使用特点、常用服务方法及应用场景,为相关人员提供了实践见解。
AI中文摘要:
大语言模型(LLM)被集成到软件系统和AI服务中,使得高效的LLM服务成为软件工程领域的关注点。为LLM提供服务颇具挑战性,因为推理过程需要消耗计算、内存、GPU资源与执行时间,同时还要维持低延迟和高吞吐量。尽管已有研究提出了LLM推理、优化及服务相关的技术与框架,但这些技术在实践中的应用情况尚不明确。本研究调查了开源软件系统中LLM服务框架和服务方法的应用情况,识别并分析了5种LLM专用框架:vLLM、SGLang、TensorRT-LLM、LMDeploy和FlashInfer。我们研究了这些框架与技术单独及组合使用的情况,不同LLM类别间的应用差异,以及代码仓库在意图、重点、用例和架构设计上的差异。结果显示,vLLM在流行度和应用范围上是最受关注的框架,而并行计算、内存管理和网络剪枝是最常用的服务方法类别。多框架使用情况有限,表明开发者依赖单一服务框架;不过,组合使用的框架可连接服务栈中的互补能力。框架应用随模型家族、模态、模型规模、领域专业化程度及部署环境而变化。代码仓库层面的分析表明,LLM服务框架支持各类应用与架构,包括基于强化学习(RL)的推理、多模态生成与理解、微服务以及云基础设施。总体而言,本研究对实践中LLM服务框架的应用进行了大规模的实证刻画,为从事LLM系统研究的人员、框架维护者和从业者提供了见解。
英文摘要:
Large Language Models (LLMs) are integrated into software systems and AI services, making efficient LLM serving a concern for software engineering. Serving LLMs is challenging because inference requires computation, memory, GPU resources, and execution while maintaining latency and throughput. Although prior research has proposed LLM inference, optimization, and serving techniques and frameworks, little is known about how they are adopted in practice. In this study, we investigate the use of LLM serving frameworks and serving methods in open-source software systems. We identify and analyze five LLM-specific frameworks: vLLM, SGLang, TensorRT-LLM, LMDeploy, and FlashInfer. We examine how these frameworks and techniques are adopted individually and in combination, how adoption varies across categories of LLMs, and how repositories differ in intent, focus, use case, and architectural design. Our results show that vLLM is the most visible framework in popularity and adoption, while parallel computation, memory management, and network pruning are the most frequently used serving-method categories. Multi-framework usage is limited, suggesting that developers rely on a single serving framework; however, combined frameworks connect complementary capabilities across the serving stack. Framework adoption varies across model families, modalities, model sizes, domain specializations, and deployment settings. Repository-level analysis shows that LLM serving frameworks support applications and architectures, including Reinforcement Learning (RL)-based reasoning, multimodal generation and understanding, microservices, and cloud infrastructure. Overall, this study provides a large-scale empirical characterization of LLM serving framework adoption in practice and offers insights for researchers, framework maintainers, and practitioners working on LLM systems.