发表机构
University of St Andrews; Zhejiang University of Technology(圣安德鲁斯大学; 浙江工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出AgBench基准套件,评估个人AI设备上本地、混合和云端智能体执行,发现本地执行成本低但成功率低,混合执行可提升成功率,部署需权衡任务成功率、成本与数据暴露。
AI 中文摘要
智能体AI系统日益依赖云端托管的大语言模型进行规划、工具使用和迭代执行,这引发了API成本和数据暴露方面的担忧。个人AI设备的进步使得智能体能够在本地执行,但设备上的有限资源可能影响任务成功率和性能。现有基准测试不足以系统地刻画这些在不同设备、工作负载和部署架构之间的权衡。我们提出了AgBench,一个基准测试套件和开放工件,用于在个人设备上对智能体AI进行可复现的评估。利用AgBench,我们评估了智能体工作负载下的本地、混合和云端执行,考察了任务成功率、延迟、云端API成本和数据暴露。我们的结果基于超过1.6207亿个数据点,表明个人AI设备能够在本地完成许多智能体任务,但仅本地执行通常比仅云端执行的任务成功率和完成时间表现更差,尤其是在并发性增加时。仅本地执行消除了云端模型API成本以及向云端智能体暴露敏感信息的风险。混合执行可以提高任务成功率,但其云端成本和数据暴露取决于智能体如何分配工作和共享信息。没有单一架构在任务成功率、有效吞吐量、云端成本和数据暴露方面表现最佳;部署选择应反映预期的工作负载和设备能力。AgBench可在以下https URL获取。
英文摘要
Agentic AI systems increasingly rely on cloud-hosted large language models for planning, tool use, and iterative execution, raising concerns about API cost and data exposure. Advances in personal AI devices enable agents to execute locally, but limited resources on device may affect task success and performance. Existing benchmarks are inadequate for systematically characterizing these trade-offs across devices, workloads, and deployment architectures. We present AgBench, a benchmark suite and open artifacts for reproducible evaluation of agentic AI on personal devices. Using AgBench, we evaluate local, hybrid, and cloud execution across agentic workloads, examining task success, latency, cloud API cost, and data exposure. Our results, drawn from over 162.07 million data points, show that personal AI devices can complete many agent tasks locally, but local-only execution generally has lower task success and longer completion times than cloud-only execution, especially as concurrency increases. Local-only execution eliminates cloud model API costs and sensitive-information exposure to cloud agents. Hybrid execution can improve task success, but its cloud cost and data exposure depend on how agents divide work and share information. No single architecture performs best across task success, goodput, cloud cost, and data exposure; deployment choices should reflect the intended workload and device capabilities. AgBench is available at https://anonymous.4open.science/r/AgBench-2777.
Comments15 pages, 12 figures, including supplementary material