arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15127cs.OScs.AIcs.DCcs.MA

从大语言模型(LLM)推理到智能体工作负载:特征分析与服务系统启示

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AgentSysBench基准套件,分析智能体工作负载与传统LLM服务的6项差异,据此开展的4项设计探索可实现延迟降低、效率提升、内存减少及冗余调用节省等效果。

中文摘要 AI 辅助

智能体应用正推动AI服务从孤立的模型推理转向长期运行的工作负载,在此类工作负载中,LLM需协调工具、环境与持久状态。然而,这类工作负载的系统行为(包括延迟、成本与瓶颈产生的环节)尚未得到充分表征,导致服务系统仍依赖针对传统推理构建的假设。本文提出AgentSysBench,这是一个基准套件与测量工具包,包含10个代表性智能体应用及统一的系统级检测工具。通过受控部署与生产跟踪,我们识别出6项区分智能体工作负载与传统LLM服务的特性:(1)执行是重量级且有状态的,10个应用中有5个的非LLM组件主导延迟,沙盒工作集内存在每个会话中峰值达28GB;(2)应用由具有异构资源亲和性的组件构成——GPU绑定的推理、内存绑定的检索、CPU绑定的沙盒,其任务延迟差异最高达32倍;(3)瓶颈会在请求、模型与部署间转移;(4)生产会话在活跃步骤之间会保持空闲状态数分钟至数小时;(5)控制平面开销(辅助LLM调用及来自工具模式与观测的上下文开销)会挤占有效计算与上下文资源;(6)来自3个应用的生产跟踪显示搜索查询与网页获取存在大量跨请求冗余,暴露出巨大的缓存机会。四项设计探索表明这些发现具有可操作性:任务感知服务可降低29%-40%的延迟,通信感知部署可提升最高4.5倍的效率,状态卸载可减少4.6倍的内存使用,工具结果缓存可消除35.2%的冗余搜索调用并节省19.3%的总搜索延迟。

英文摘要

Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads---where latency, cost, and bottlenecks arise---remains poorly characterized, leaving serving systems to rely on assumptions built for conventional inference. We present AgentSysBench, a benchmark suite and measurement toolkit with ten representative agentic applications and unified systems-level instrumentation. Across controlled deployments and production traces, we identify six properties that distinguish agentic workloads from conventional LLM serving: (1) execution is heavyweight and stateful, with non-LLM components dominating latency in 5 of 10 applications and sandbox working-set memory peaking at 28 GB per session; (2) applications compose components with heterogeneous resource affinity---GPU-bound inference, memory-bound retrieval, CPU-bound sandboxes---whose task latencies diverge by up to 32x; (3) bottlenecks shift across requests, models, and deployments; (4) production sessions hold state idle for minutes to hours between active steps; (5) a control-plane tax---auxiliary LLM calls and context overhead from tool schemas and observations---crowds out productive compute and context; and (6) production traces from three applications reveal heavy cross-request redundancy in search queries and web fetches, exposing a large caching opportunity. Four design explorations demonstrate that these findings are actionable: task-aware serving reduces latency by 29--40%, communication-aware placement by up to 4.5x, state offloading reduces memory usage by 4.6x, and tool-result caching removes 35.2% of redundant search calls and saves 19.3% of aggregate search latency.

发表机构

  • Hong Kong University of Science and Technology(香港科技大学)
  • Alibaba Group(阿里巴巴集团)
  • Bytedance(字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

↑