arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24509cs.AIcs.SE

PeakBench:面向大语言模型智能体的资源感知工具调用基准测试

PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents

Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li, De-Chuan Zhan, Han-Jia Ye

首次发表
浏览论文内容

中文总结 AI 辅助

PeakBench是面向LLM智能体的资源感知工具调用基准,通过两部分评估框架解耦逻辑规划与物理调度,实验表明资源信息可减少溢出、提升利用率,为相关研究提供测试平台。

中文摘要 AI 辅助

大语言模型(LLM)智能体越来越多地通过调用多种工具来解决任务,其中并行执行对于降低延迟至关重要,但安全管理难度较大。现有的智能体基准测试主要在大多为串行执行的场景下评估工具选择、参数生成和端到端成功率,很大程度上忽略了有效的并行化和资源受限调度。这一缺失的调度维度会导致实际故障模式:串行执行安全但速度慢,而不考虑资源的并行执行速度快但容易出现可避免的资源溢出。为解决这一差距,我们推出了PeakBench,这是一个可执行多工具工作流的基准测试,带有基于执行的依赖关系注释和已测量的资源概况。评估此类工作流的核心挑战是归因:故障和低效率可能源于不正确的依赖规划、资源受限调度不佳,或两者兼而有之。PeakBench通过一个由两部分组成的评估框架解决了这一挑战,该框架将逻辑规划与物理调度解耦,并为每个维度提供专用指标。使用该框架,我们表明在资源约束下,强大的逻辑规划并不能可靠地转化为安全或高效的执行。我们进一步表明,暴露资源信息可以减少可避免的溢出并提高资源利用率,这使得PeakBench成为诊断资源感知智能体行为的有用测试平台。代码可在此https URL获取。

英文摘要

LLM agents increasingly solve tasks by invoking multiple tools, where parallel execution is essential for low latency but difficult to manage safely. Existing agent benchmarks primarily evaluate tool selection, argument generation, and end-to-end success under mostly serial execution, largely overlooking valid parallelization and resource-constrained scheduling. This missing scheduling dimension creates a practical failure mode: serial execution is safe but slow, while resource-agnostic parallel execution is fast but prone to avoidable resource overflows. To address this gap, we introduce PeakBench, a benchmark of executable multi-tool workflows with execution-grounded dependency annotations and measured resource profiles. A central challenge in evaluating such workflows is attribution: failures and inefficiencies may arise from incorrect dependency planning, poor resource-constrained scheduling, or both. PeakBench addresses this challenge with a two-part evaluation framework that disentangles logical planning from physical scheduling, with dedicated metrics for each dimension. Using this framework, we show that strong logical planning does not reliably translate into safe or efficient execution under resource constraints. We further show that exposing resource information can reduce avoidable overflows and improve resource utilization, making PeakBench a useful testbed for diagnosing resource-aware agent behavior. Code is available at https://github.com/Czzzk/Staggering-the-Peaks.

发表机构

  • School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)
  • Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑