arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ClawProBench:基于运行时覆盖与冻结工作空间保留集的轨迹感知AI智能体评估

ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts

YuanHang Xiao

arXiv 2608.22510首次发表:更新:

发表机构

The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出ClawProBench基准,基于OpenClaw运行时构建,通过含102场景的全配置集与68场景的冻结保留集评估AI智能体,采用安全门控公式评分,发现最终答案排行榜存在缺陷,不同评估视角的智能体排名差异显著。

AI 中文摘要

智能体基准测试通常仅评估最终答案,即便智能体在有状态运行时上运行。我们认为这会导致评估的规格不足:正确的评估单元是声明的模型-运行时配置,其故障可能出现在证据获取、运行时路由、安全边界或重复执行环节。我们提出ClawProBench,这是一个针对运行时原生智能体评估的轨迹感知基准,基于OpenClaw实例化——OpenClaw是一个具备工作空间工具及原生浏览、记忆、消息、调度、技能、子智能体界面的实时智能体运行时。ClawProBench定义了两个赛道:包含实时工作空间与原生运行时路由任务的102场景全配置集,以及具备封闭世界JSON输出契约用于稳健排名的68场景冻结保留集。试验通过安全门控公式从执行轨迹中评分,该公式结合了正确性、过程质量与效率,并保留故障证据以供审计。我们的匿名制品包含基准定义、评分代码、清单及净化后的轨迹。我们在全配置集上评估68种配置,在保留集上评估37种配置。最高安全门控平均轨迹分数为0.7671;原生运行时任务表现逊于工作空间实时任务(0.5238 vs 0.6415);在保留集上,pass@k-any优于严格三次试验通过(0.6638 vs 0.2890),而全配置集与保留集的排名显示弱一致性(斯皮尔曼相关系数0.1300);仅基于正确性的排名与过程感知、安全门控及严格通过的排名差异显著;最终答案排行榜可能隐藏原生界面弱点、一次性成功及轨迹局部智能体故障模式。

英文摘要

Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-specifies what is being evaluated: the proper unit is a declared model-plus-runtime configuration whose failures can occur in evidence acquisition, runtime routing, safety boundaries, or repeated execution. We present ClawProBench, a trace-aware benchmark for runtime-native agent evaluation instantiated on OpenClaw, a live agent runtime with workspace tools and native surfaces for browsing, memory, messaging, scheduling, skills, and subagents. ClawProBench defines two tracks: a 102-scenario full profile with live workspace and native-runtime routing tasks, and a frozen 68-scenario holdout with closed-world JSON output contracts for robust ranking. Trials are scored from execution traces via a safety-gated formula combining correctness, process quality, and efficiency, preserving failure evidence for audit. Our anonymous artifact includes benchmark definitions, scoring code, manifests and sanitized traces. We evaluate 68 configurations on the full profile and 37 on holdout. The top safety-gated average trace score is 0.7671. Native-runtime tasks underperform workspace-live tasks (0.5238 vs. 0.6415). On holdout, pass@k-any outperforms strict three-trial pass (0.6638 vs. 0.2890), while full-profile and holdout rankings show weak alignment (Spearman 0.1300). Rankings based purely on correctness differ substantially from process-aware, safety-gated and strict-pass views. Final-answer leaderboards may hide native-surface weaknesses, one-off successes and trace-local agent failure modes.

Comments29 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑