发表机构
Microsoft Azure(微软Azure)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过生产与受控研究分析智能体AI工作流的架构特性,揭示传统服务器的不匹配问题,提出Agora原型优化资源利用与吞吐量,为未来服务器架构指明方向。
AI 中文摘要
智能体AI正在数据中心兴起,但其架构影响仍未被探索。我们对智能体工作流进行分类,并通过在Microsoft Azure开展的生产研究以及对开源框架的受控研究,首次对其进行了架构表征。研究表明,智能体执行是碎片化且异构的,请求会扩展为包含LLM推理、工具调用和编排决策的工作流,这些操作会反复跨越CPU-GPU边界。我们的分类解释了这种碎片化如何转化为资源需求:由于编排和工具在主机上运行,CPU处于关键路径;执行结构决定了负载随时间的变化,其负载保持较低但会出现突发峰值;模型组合决定了工作流对GPU的利用均匀程度,任务和工具的多样性会进一步扩大该范围。这些特性暴露了传统统一服务器的架构不匹配问题:碎片化的执行会在突发需求下浪费CPU和GPU容量;不同的软件角色使统一CPU配置效率低下;将多个智能体复用至共享核心会降低微架构局部性。基于上述发现,我们推导了智能体服务器的架构影响,并通过我们的商用服务器原型Agora进行了验证。Agora动态收集闲置CPU核心用于共置吞吐量工作,同时保护智能体的尾部延迟免受工具峰值影响;它通过在每个GPU上放置更多智能体、预取下一个智能体的状态以隐藏交换延迟来超额分配GPU内存;为使机器适配异构角色,Agora按角色对核心进行池化,并应用感知亲和性的调度来恢复局部性,还会自动调整机制以适配工作负载。Agora在保持智能体尾部延迟的同时,提高了利用率和服务器吞吐量,我们的见解也为智能体AI的未来服务器架构指明了关键方向。
英文摘要
Agentic AI is emerging in datacenters, but its architectural implications remain unexplored. We organize agentic workflows in a taxonomy and present its first architectural characterization with a production study at Microsoft Azure and a controlled study of open-source frameworks. We show that agentic execution is fragmented and heterogeneous. Requests expand into a workflow of LLM inferences, tool invocations, and orchestration decisions that repeatedly cross the CPU-GPU boundary. Our taxonomy explains how this fragmentation turns into resource demand. As orchestration and tools run on the host, the CPU sits on the critical path. Execution structure sets the load over time, which stays low with sudden spikes. Model composition sets how evenly the workflow uses the GPUs. Diversity in tasks and tools widens this range even further. These characteristics expose architectural mismatches of conventional uniform servers. Fragmented execution strands CPU and GPU capacity despite bursty demand. Different software roles make homogeneous CPU provisioning inefficient. Finally, multiplexing many agents onto shared cores degrades microarchitectural locality. Guided by our findings, we derive implications for agentic servers and examine them through Agora, our prototype for commodity servers. Agora dynamically harvests idle CPU cores for co-located throughput work, while protecting agentic tail latency against tool spikes. It oversubscribes GPU memory by placing more agents on each GPU, prefetching the next agent's state to hide swap latency. To match the machine to the heterogeneous roles, Agora pools cores by role and applies affinity-aware scheduling to restore locality. It automatically tunes mechanisms to the workload. Agora improves utilization and server throughput while preserving agent tail latency. Our insights also identify key directions for future server architectures for agentic AI.