发表机构
Imperial College London; University of Cambridge; University of Oxford(帝国理工学院; 剑桥大学; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AgentPerfBench提出面向智能体推理的基准套件,利用真实轨迹与合成配置,揭示现有基准忽略上下文增长和硬件饱和的问题,并通过NCU轨迹和roofline模型量化聊天到智能体的性能差距。
AI 中文摘要
LLM服务引擎(如vLLM和SGLang)的优化在很大程度上是基准驱动的:优化方案、调度策略、硬件和系统设计均基于代表性工作负载进行选择。然而,在智能体时代出现了显著的不匹配。现有基准主要关注简单的单轮聊天工作负载。LLM应用日益智能体化:编码智能体、终端执行系统和工具使用智能体发出多轮请求,且上下文长度不断增长。我们推出AgentPerfBench,一个面向智能体推理的基准套件。它使用来自智能体基准(如SWE-Bench和TerminalBench)的真实轨迹,以及标准聊天基线。这使得能够在涉及工具调用、技能利用和上下文长度不断增长的多轮任务上对模型进行基准测试。AgentPerfBench还从真实轨迹中推导出的输入长度、输出长度和轮次数的经验分布中采样,生成具有代表性的合成配置文件,用于在新硬件上进行廉价且准确的测量。此外,我们进一步发现,现有若干基准因两个关键原因未能准确反映真实硬件性能:1)它们未考虑现实中的上下文长度增长;2)它们在未达到硬件饱和的情况下测量推理性能。我们详细讨论这些问题,并提供丰富的内核级Nsight Compute(NCU)轨迹,以构建一个新的多维roofline模型,该模型捕捉内存带宽和内存容量占用方面的硬件-系统限制。该基准套件随后包含自动化脚本,以在使用多样化智能体轨迹评估时,识别新兴硬件上的潜在瓶颈条件。综合而言,这些贡献量化了当前推理基准中聊天到智能体的差距,并通过roofline分析表征了每个内核的GPU资源利用率。
英文摘要
The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system designs are all selected based on representative workloads. However, a significant mismatch has emerged in the agentic era. Existing benchmarks primarily focus on simple single-turn chatbot workloads. LLM applications are increasingly agentic: coding agents, terminal execution systems, and tool-use agents issue multi-turn requests with growing context lengths. We introduce AgentPerfBench, a benchmark suite for agentic inference. It uses real traces from agentic benchmarks, such as SWE-Bench and TerminalBench, alongside standard chat baselines. This enables benchmarking of models on multi-turn tasks involving tool calling, skill utilization, and increasing context lengths. AgentPerfBench also samples from empirical distributions of input length, output length, and turn count derived from the real traces, generating representative synthetic profiles for cheap and accurate measurements on new hardware. In addition, we further find that several existing benchmarks fail to accurately reflect real hardware performance for two key reasons: 1) they do not account for realistic context-length growth, and 2) they measure inference performance without operating at hardware saturation. We discuss these issues in detail and provide rich kernel-level Nsight Compute (NCU) traces to construct a new multi-dimensional roofline model that captures hardware-system limitations in both memory bandwidth and memory capacity footprint. The benchmarking suite then includes automated scripts to identify potential bottleneck conditions on emerging hardware when evaluated with diverse agentic traces. Together, these contributions quantify the chat-to-agentic gap in current inference benchmarks and characterise per-kernel GPU resource utilisation via roofline analysis.