DAREBench:面向部署的模型作为智能体的可靠评估
DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
- Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
- School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
- MiLM Plus, Xiaomi Inc.(小米公司MiLM Plus)
- Department of Computer Science, Brown University(布朗大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
DAREBench提出面向部署的可靠智能体评估基准,基于OpenClaw环境组织233个任务,评估35个模型,发现无单一模型全能,部署需考虑工作负载与成本权衡。
AI中文摘要:
随着大型语言模型从问答系统演变为通用智能体,评估必须超越静态答案正确性,以评估多模态感知、多步骤执行、工具使用和工件交付。然而,现有基准通常与特定任务类型、执行环境或评分协议绑定,限制了其在部署决策中的可比性、可解释性和可靠性。我们提出DAREBench(面向部署的模型作为智能体的可靠评估),一个旨在捕捉工作负载变化并支持可靠智能体评估的基准。基于共享的OpenClaw执行环境,DAREBench将选自22个源基准的233个任务组织成一个由输入模态和执行形式定义的$2\ imes3$工作负载矩阵,并在统一的基于契约的协议下进行评估,辅以基于证据的评分审计。我们评估了23个商业API模型和12个本地部署的开源权重模型,共进行7,587次模型-任务运行,报告了准确性和令牌消耗以及API模型的参考成本。结果表明,没有单一模型在所有工作负载组中占主导地位,文本和多模态任务表现出不同的准确性-成本权衡,本地开源权重模型在若干组中具有竞争力,但总体上仍落后于前沿商业模型。这些发现表明,智能体部署和模型选择应考虑工作负载概况、部署模式以及准确性-成本权衡,而非依赖单一的总体得分。
英文摘要:
As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are often tied to specific task types, execution environments, or scoring protocols, limiting their comparability, interpretability, and reliability for deployment decisions. We introduce DAREBench (Deployment-Aware and Reliable Evaluation of Models as Agents), a benchmark designed to capture workload variation and support reliable agent evaluation. Built on a shared OpenClaw execution environment, DAREBench organizes 233 tasks selected and adapted from 22 source benchmarks into a $2\times3$ workload matrix defined by input modality and execution form, and evaluates them under a unified contract-based protocol with evidence-based score auditing. We evaluate 23 commercial API models and 12 locally deployed open-weight models over 7,587 model--task runs, reporting accuracy and token consumption alongside reference costs for API models. Results show that no single model dominates all workload groups, text and multimodal tasks exhibit distinct accuracy--cost trade-offs, and local open-weight models are competitive in several groups but still trail frontier commercial models overall. These findings suggest that agent deployment and model selection should consider workload profiles, deployment mode, and accuracy--cost trade-offs rather than rely on a single aggregate score.