arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13573cs.AI

LLM 服务的一年:工作负载演变、缓存与负载均衡

A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing

William Nixon, Jon Durbin, Florian Standhartinger, Haryadi S. Gunawi, Juncheng Yang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过分析 Chutes 一年的 LLM 生产请求轨迹,揭示了 LLM 服务工作负载的演变规律与用户-模型结构,并将发布完整轨迹供后续研究使用。

中文摘要 AI 辅助

大型语言模型(LLM)服务已成为关键的云工作负载,而真实的请求轨迹对于驱动和基准测试服务系统至关重要。然而,现有的 LLM 服务工作负载研究在规模和范围上仍然有限,它们通常仅观测短时间周期,且对生产环境中用户与模型的交互情况可见性有限,因此无法完全捕捉 LLM 服务工作负载随时间的演变情况,也无法完全捕捉用户-模型交互如何塑造生产流量。在本研究中,我们通过对 Chutes 提供的一年生产请求轨迹进行全局特征分析和纵向研究,加深了对真实世界 LLM 服务工作负载的理解。与先前的研究不同,我们的轨迹捕获了众多模型和用户的完整生产行为,包括热门模型和长尾模型。我们从聚合、时间、模型级别和用户级别四个视角分析工作负载,揭示了通常被聚合视图隐藏的工作负载演变和用户-模型结构。为支持未来研究,我们将随本文发布完整的一年请求轨迹,使下游研究无需依赖采样或人工生成的工作负载即可开展生产行为研究。

英文摘要

Large Language Model (LLM) serving has become a critical cloud workload, and realistic traces are essential for motivating and benchmarking serving systems. However, existing LLM serving workload studies remain limited in scale and scope. They often observe short time periods and provide limited visibility into how users interact with models in production. As a result, they do not fully capture how LLM serving workloads evolve over time or how user-model interactions shape production traffic. In this work, we further the understanding of real-world LLM serving workloads through both a global characterization and a longitudinal study of a one-year production trace from CompanyX. Unlike prior studies, our trace captures full production behavior across many models and users, including both popular and long-tail models. We analyze the workload from aggregate, temporal, model-level, and user-level perspectives, revealing workload evolution and user-model structure that are typically hidden behind aggregate views. To support future research, we publicly release the full one-year trace, enabling downstream studies of production behavior without relying on sampled or synthetically generated workloads. The trace is available at https://github.com/HarvardMadSys/chutes_workload.

发表机构

  • University of Chicago(芝加哥大学)
  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

↑