AI 中文总结
本研究提出AgentSLABench框架,在资源约束下评估自主AI智能体,通过多维度指标分析发现专用智能体表现优于通用基线,验证了效率调整成功率的重要性并开放相关资源。
AI 中文摘要
我们提出了AgentSLABench,这是一种面向自主AI智能体的资源感知评估框架,该框架在声明的资源预算下,除了测量正确性外,还测量延迟、成本、计算资源、内存和网络使用情况。与仅报告准确率的标准基准不同,AgentSLABench为每个智能体、每个任务生成多维度性能轮廓——其方式与系统分析工具(perf、pprof、cProfile)测量代码资源消耗类似,但将任务正确性作为首要维度纳入考量。AgentSLABench提供6个类别共16个任务环境(其中核心任务5项:多跳问答、零售替代、代码生成、网络购物、旅行规划;扩展任务11项),配备隔离的Docker容器、声明式CPU/内存/时间/网络预算、带有SHA256哈希的密封测试集,以及标准化的分析协议。我们对5个通用基线智能体(ReAct、PlanAndSolve、Reflexion、CoT、Random)和4个任务专用智能体进行了分析,发现专用智能体在5项核心任务中的3项(fact_qa、web_shopping、travel_planning)实现了100%的成功率,在零售和代码生成任务上的成功率为66.7%-83.3%,而通用基线智能体在4项领域任务上完全失败。关键的是,我们报告了效率调整成功率(EASR)——即相对于声明预算、按资源消耗加权的成功率,这表明无限制成本下的高准确率并不具备生产可行性。我们发布了完整的基础设施、密封测试集和分析结果,以支持可复现的、资源感知的智能体评估。
英文摘要
We present AgentSLABench, a resource-aware evaluation framework for autonomous AI agents that measures correctness alongside latency, cost, compute, memory, and network usage under declared resource budgets. Unlike standard benchmarks that report only accuracy, AgentSLABench produces a multi-dimensional profile per agent per task - the same way systems profilers (perf, pprof, cProfile) measure resource consumption of code, but extended with task correctness as a first-class dimension. AgentSLABench provides 16 task environments across 6 categories (5 core: multi-hop QA, retail substitution, code generation, web shopping, travel planning; 11 extended) with isolated Docker containers, declared CPU/memory/time/network budgets, sealed test sets with SHA256 hashes, and a standardized profiling protocol. We profile 5 general-purpose baseline agents (ReAct, PlanAndSolve, Reflexion, CoT, Random) plus 4 task-specialized agents, finding that specialized agents achieve 100% success on 3/5 core tasks (fact_qa, web_shopping, travel_planning) and 66.7-83.3% on retail and code_gen, while general baselines fail entirely on 4/5 domain tasks. Crucially, we report the Efficiency-Adjusted Success Rate (EASR) - success weighted by resource consumption relative to declared budgets - revealing that high accuracy at unbounded cost is not production-viable. We release the full infrastructure, sealed test sets, and profiling results to enable reproducible, resource-aware agent evaluation.
Comments7 pages, 2 figures, 9 tables. Code, sealed test sets, and profiling artifacts available at: https://github.com/MeherBhaskar/agentslabench