从大语言模型(LLM)推理到智能体工作负载:特征分析与服务系统启示
From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems
浏览论文内容
中文总结 AI 辅助
本文提出AgentSysBench基准套件,分析智能体工作负载与传统LLM服务的6项差异,据此开展的4项设计探索可实现延迟降低、效率提升、内存减少及冗余调用节省等效果。
中文摘要 AI 辅助
智能体应用正推动AI服务从孤立的模型推理转向长期运行的工作负载,在此类工作负载中,LLM需协调工具、环境与持久状态。然而,这类工作负载的系统行为(包括延迟、成本与瓶颈产生的环节)尚未得到充分表征,导致服务系统仍依赖针对传统推理构建的假设。本文提出AgentSysBench,这是一个基准套件与测量工具包,包含10个代表性智能体应用及统一的系统级检测工具。通过受控部署与生产跟踪,我们识别出6项区分智能体工作负载与传统LLM服务的特性:(1)执行是重量级且有状态的,10个应用中有5个的非LLM组件主导延迟,沙盒工作集内存在每个会话中峰值达28GB;(2)应用由具有异构资源亲和性的组件构成——GPU绑定的推理、内存绑定的检索、CPU绑定的沙盒,其任务延迟差异最高达32倍;(3)瓶颈会在请求、模型与部署间转移;(4)生产会话在活跃步骤之间会保持空闲状态数分钟至数小时;(5)控制平面开销(辅助LLM调用及来自工具模式与观测的上下文开销)会挤占有效计算与上下文资源;(6)来自3个应用的生产跟踪显示搜索查询与网页获取存在大量跨请求冗余,暴露出巨大的缓存机会。四项设计探索表明这些发现具有可操作性:任务感知服务可降低29%-40%的延迟,通信感知部署可提升最高4.5倍的效率,状态卸载可减少4.6倍的内存使用,工具结果缓存可消除35.2%的冗余搜索调用并节省19.3%的总搜索延迟。
英文摘要
Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads---where latency, cost, and bottlenecks arise---remains poorly characterized, leaving serving systems to rely on assumptions built for conventional inference. We present AgentSysBench, a benchmark suite and measurement toolkit with ten representative agentic applications and unified systems-level instrumentation. Across controlled deployments and production traces, we identify six properties that distinguish agentic workloads from conventional LLM serving: (1) execution is heavyweight and stateful, with non-LLM components dominating latency in 5 of 10 applications and sandbox working-set memory peaking at 28 GB per session; (2) applications compose components with heterogeneous resource affinity---GPU-bound inference, memory-bound retrieval, CPU-bound sandboxes---whose task latencies diverge by up to 32x; (3) bottlenecks shift across requests, models, and deployments; (4) production sessions hold state idle for minutes to hours between active steps; (5) a control-plane tax---auxiliary LLM calls and context overhead from tool schemas and observations---crowds out productive compute and context; and (6) production traces from three applications reveal heavy cross-request redundancy in search queries and web fetches, exposing a large caching opportunity. Four design explorations demonstrate that these findings are actionable: task-aware serving reduces latency by 29--40%, communication-aware placement by up to 4.5x, state offloading reduces memory usage by 4.6x, and tool-result caching removes 35.2% of redundant search calls and saves 19.3% of aggregate search latency.
发表机构
- Hong Kong University of Science and Technology(香港科技大学)
- Alibaba Group(阿里巴巴集团)
- Bytedance(字节跳动)
机构由 AI 辅助整理,请以论文原文为准。