arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非所有AI智能体都相同:资源与性能动态特征分析

Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

Wonmi Choi, Minuk Park, Zhixiong Niu, Yongqiang Xiong, Chuck Yoo, Gyeongsik Yang

arXiv 2609.19947首次发表:更新:

发表机构

Korea University; Microsoft Research Asia(高丽大学; 微软亚洲研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过分析检索增强问答、网络搜索和软件编码三类任务的资源动态,发现智能体资源行为因任务而异,并提出CPU感知工具准入与任务感知CPU分配,使CPU敏感任务延迟提升5.4倍、平均延迟降低32%。

AI 中文摘要

基于大语言模型(LLM)的AI智能体通过迭代推理和工具执行来处理用户请求,通常涉及调用远程LLM API并配合本地工具容器。这种执行模式使得智能体服务的优化变得困难,因为延迟、本地资源需求和容器瓶颈在请求之间相互交织。然而,当前的智能体生态系统在运行时很少考虑资源动态,导致宝贵资源的严重浪费。本文针对三个代表性任务分析了AI智能体的资源混合特征:检索增强问答、网络搜索和软件编码。为此,我们根据并发处理多个请求和任务的资源动态来刻画延迟特征。我们的测量表明,智能体根据任务不同表现出广泛的行为差异,以至于即使是相同的工具,其资源动态也可能存在显著不同。我们还发现,并发运行多个请求会暴露任务相关的资源动态瓶颈,如CPU、磁盘I/O和内存。此外,我们发现更快的LLM响应或更多的CPU核心并不总能加速智能体。基于这些观察,我们展示了利用任务资源动态的新优化机会:CPU感知的工具准入和任务感知的CPU分配。结果表明,CPU敏感型智能体任务的延迟提升了约5.4倍,多个任务的平均延迟相比原生智能体降低了约32%。

英文摘要

LLM-based AI agents process user requests through iterative reasoning and tool execution, often involving the invocation of remote LLM APIs with local tool containers. This execution model can make the optimization of agent serving difficult because latency, local resource demand, and container bottlenecks inter-mix across requests. However, the current agent ecosystem runs without much consideration of resource dynamics, which results in significant waste of the precious resources. This paper analyzes the resource inter-mix of AI agents for three representative tasks: retrieval-augmented question answering, web search, and software coding. To this end, we characterize the latency with respect to the resource dynamics of processing multiple requests and tasks concurrently. Our measurements show that agents have a wide range of behaviors depending on tasks, so that even the same tool can differ substantially in resource dynamics. We also find that running multiple requests concurrently exposes task-dependent bottlenecks in resource dynamics such as CPU, disk I/O, and memory. Furthermore, we uncover that faster LLM responses or more CPU cores do not always accelerate agents. Based on these observations, we demonstrate new optimization opportunities that exploit the resource dynamics of tasks: CPU-aware tool admission and task-aware CPU allocation. Our results show that the latency of CPU-sensitive agent tasks improves $\sim$5.4$\times$, and the average latency across multiple tasks is reduced $\sim$32% compared to native agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑