arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FineServe:一个细粒度的数据集以及对全球大语言模型服务工作负载的特征描述

FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads

Tiancheng Zhang, Shaoyuan Huang, Mingyuan Wang, Yunfeng Zhao, Xiaofei Wang, Wenyu Wang

arXiv 2607.19349首次发表:更新:

发表机构

Tianjin University; PPIO Cloud (Shanghai) Co., Ltd.(天津大学; PPIO云(上海)有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型服务中高效服务面临的挑战,提出FineServe数据集,能对异构模型和任务的服务动态进行细粒度特征描述。通过分析得出波动模式,开发工作负载生成器,为评估相关策略提供现实基础。

AI 中文摘要

大语言模型(LLMs)越来越多地作为随时可用的在线服务进行部署,使得高效的大语言模型服务成为一个关键的系统挑战。在多变的需求下实现低延迟和高吞吐量需要深入了解实际的服务工作负载,然而现有研究往往依赖代理跟踪或粗粒度特征描述,无法捕捉现代多模型大语言模型平台的异质性。我们展示了FineServe,这是一个从全球商业市场收集的多模型大语言模型服务工作负载数据集,能够对跨异构模型和任务的实际服务动态进行细粒度特征描述。利用FineServe,我们对到达动态和令牌行为进行了全面分析,揭示了跨模型架构、规模和任务意图的根本不同的波动模式。基于这些见解,我们开发了FineServe工作负载生成器,它将细粒度的模型感知工作负载组合成可配置的混合,专为基准测试多模型服务平台量身定制。通过揭示这些细粒度的工作负载动态,FineServe为评估大语言模型服务系统中的路由、调度和容量规划策略提供了一个现实的基础。FineServe可在这个https网址获取。

英文摘要

Large language models (LLMs) are increasingly deployed as always-on online services, making efficient LLM serving a critical systems challenge. Achieving low latency and high throughput under volatile demand requires deep understanding of real-world serving workloads, yet existing studies often rely on proxy traces or coarse-grained characterizations that fail to capture the heterogeneity of modern multi-model LLM platforms. We present FineServe, an in-the-wild, multi-model LLM serving workload dataset collected from a global commercial marketplace, enabling fine-grained characterization of real-world serving dynamics across heterogeneous models and tasks. Leveraging FineServe, we conduct a comprehensive analysis of arrival dynamics and token behavior, revealing fundamentally different fluctuation regimes across model architectures, scales and task intents. Building on these insights, we develop the FineServe workload generator, which composes fine-grained model-aware workloads into configurable mixtures tailored for benchmarking multi-model serving platforms. By exposing these fine-grained workload dynamics, FineServe provides a realistic foundation for evaluating routing, scheduling, and capacity-planning strategies in LLM serving systems. FineServe is available at https://github.com/hihiztc1/FineServe.

Comments14 pages, 14 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑