arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20370cs.DCcs.MA

使用XPerf对智能体AI工作负载的大语言模型服务系统进行基准测试

Benchmarking LLM Serving Systems for Agentic AI Workloads with XPerf

  • IBM Research(IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

Michael Wang, Yikang Yue, Shaobo Li, Yirui Eric Zhou, Chen Wang, Jian Huang

AI总结:

研究人员提出XPerf基准测试框架,可对智能体AI工作负载的LLM服务系统进行负载测试,能减少工作负载变化、提供性能分析并助力服务调试,将在GitHub开源。

AI中文摘要:

我们提出了XPerf,这是一个对大语言模型(LLM)服务系统进行负载测试的基准测试框架,适用于各类智能体AI工作负载。它能对服务系统和硬件进行详细分析,使用户可识别智能体工作负载引入的性能瓶颈。在智能体工作负载下对LLM服务系统进行基准测试颇具挑战——智能体应用依赖非确定性的LLM输出来指导控制流,因此每次运行的工作负载模式存在不可预测的变化。XPerf通过细粒度的轨迹重放方法最大程度减少这种工作负载变化:它使用户能轻松从真实智能体应用收集轨迹,按需合成各种模式的新工作负载,并在不同LLM服务系统上可重复地重放这些工作负载。XPerf默认包含8个不同用例的智能体应用(如编码、深度研究和问答)。我们使用这些工作负载的实证研究表明,XPerf可准确重放智能体工作负载、提供详细的性能分解、扩展至更大规模的服务系统,并助力服务系统调试。我们将在GitHub上开源XPerf。

英文摘要:

We present XPerf, a benchmarking framework that load-tests LLM serving systems with diverse agentic AI workloads. It provides detailed profiling of the serving system and hardware, enabling users to identify performance bottlenecks introduced by agentic workloads. Benchmarking LLM serving systems under agentic workloads is challenging - agentic applications rely on nondeterministic LLM outputs to guide their control flow; therefore, workload patterns vary unpredictably from run to run. XPerf minimizes this workload variation with a fine-grained trace replay approach: it enables users to easily collect traces from real agentic applications, synthesize new workloads with various patterns if needed, and reproducibly replay them on different LLM serving systems. XPerf includes eight agentic applications across diverse use cases (e.g., coding, deep research, and Q&A) by default. Our empirical study using these workloads shows that XPerf accurately replays agentic workloads, provides detailed performance breakdowns, scales to larger serving systems, and assists in serving system debugging. We will open-source XPerf on GitHub.

↑