arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22808cs.LGcs.MAcs.PF

CatchBench:智能体故障何时能被检测?

CatchBench: When Can an Agent Failure Be Caught?

Yue Zhao, Mengyuan Li, Ruolin Li, Prince Zizhuang Wang, Shuli Jiang, Linsey Pang, Xiongye Xiao, Xiyang Hu

首次发表
浏览论文内容

中文总结 AI 辅助

CatchBench是首个在同一任务-方法接口下对智能体运行前、运行中、完成后三种信息状态进行评分的基准,涵盖多类模型与配置,揭示基准分数需结合标签过程才可解释。

中文摘要 AI 辅助

智能体故障何时能被检测?审计通常受限于记录而非方法。因此CatchBench将审计员的问题置于三种信息状态下:运行前的声明配置(PRE)、不断增长的轨迹前缀(LIVE)以及完成的轨迹(POST)。现有基准通常固定其中一种状态或改变遥测方式;据我们所知,没有任何基准能在同一任务-方法接口下对这三种状态进行评分。每种状态对应不同问题,因此7项任务合约带有各自的标签和指标,而非单一排行榜。其中4项为证据类,3项为Gold衍生的机制诊断类。本次发布对72个参赛项目进行了评分,涵盖规则扫描器、结构模型以及9个模型家族中的11个LLM评判器(包括GPT、Claude、Gemini、Gemma、Llama、Qwen、DeepSeek、Mistral、Nova),涉及1187个声明配置和1162条记录运行。多数竞技场结果无法排序:118个预先声明的对比中有47个可区分,其余未解决的结果会被公布而非排名。两项最明确的结果与我们自身的数据相悖。其中一项规则忽略所有名称和权限,仅标记首次声明后出现的每项能力;在6个配置来源中的一个上,它达到了完美的F1分数,因此该分数衡量的是语料库的构建方式,而非方法的推理能力。随后我们的可接受性标准拒绝了一个注入的子strate,并拒绝了另一个的证据状态。因此,基准分数只有在其标签背后的过程被公布并针对可能存在的捷径进行测试后才具有可解释性。我们同时报告了这些内容,并使用发布的预测重新生成了所有排序,未调用任何模型。

英文摘要

When can an agent failure be caught? A weak audit score alone cannot identify whether the record or the method is limiting. CatchBench therefore puts one auditor's question to three information states: the declared configuration before a run (PRE), a growing prefix of its trace (LIVE), and the finished trace (POST). Prior benchmarks fix one of these states or vary the telemetry; to our knowledge none scores all three under one task-method interface. Each state admits different questions, so seven task contracts carry their own labels and metrics rather than one leaderboard. Four are evidential; three are Gold-derived mechanism diagnostics. The release scores 72 entrants, from rule scanners and structural models to eleven LLM judges across nine model families (GPT, Claude, Gemini, Gemma, Llama, Qwen, DeepSeek, Mistral, Nova), over 1187 declared configurations and 1162 recorded runs. Every recorded comparison is published as a measured difference with its interval, uncorrected, and no board declares a winner it cannot show. The three sharpest results cut against our own data. One rule reads declaration order alone and reaches a perfect F1 on one of six configuration sources, so a score there measures how the corpus was built. Our admissibility bar then rejected one injected substrate and withheld evidential status from the other. A published structural gain also turns on which size reference it is measured against. A benchmark number is therefore not interpretable until the process behind its labels is published and tested for the shortcut it may leave. We report all three, and regenerate every ordering from released predictions with no model call.

发表机构

  • University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑