arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37315cs.SEcs.AIcs.LG

智能体基准测试是否名副其实?对使用工具的智能体环境的可执行契约审计

Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments

发表机构佛罗里达国际大学
查看机构详情
  • Florida International University(佛罗里达国际大学)

机构由 AI 辅助整理,请以论文原文为准。

Rohith Reddy Bellibatlu, Zichong Wang, Wenbin Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过可执行契约审计四个工具使用智能体基准测试,发现七个工具缺陷及评估器属性问题,表明基准分数可能反映接口宣称而非实际状态变化。

中文摘要 AI 辅助

使用工具的智能体正进入错误行动会带来实际成本的场景,而认证这些智能体的基准测试根据每次模拟工具调用所报告的结果进行评分,假设工具执行了其接口所宣称的功能。我们调查的审计分类法没有为这一假设设置类别,而分数之下的缺陷在每次重跑时都会出现。我们将工具宣称的表面视为可执行契约,对照契约检查实现,并通过任务文件和评估器代码追踪每个分数的来源,直至源自一个有缺陷的工具本应写入的状态的判定结果。在四个基准测试的34个被审计的变更工具中,我们确认了七个工具缺陷和一个评估器属性(在固定提交版本上)。在注入的缺陷上,检查器在25个标记中未产生误报,标记了5个阴性对照中的2个,并漏掉了大多数:在33个有评分的漏报中,有29个是条款覆盖了缺陷但没有探针揭示它。检查器自身的静态部分单独运行时,标记了17个确认位点中的14个,因此在这些发现上,动态部分确认和追踪而非发现。另外12个AgentDojo工具(其中六个为留出工具,七个为首先审计的工具)完成了其25个工具的变更表面,在该表面上,按照我们的契约解读,至少有5个工具与其宣称的表面存在偏差,这一比例仅针对AgentDojo。没有黄金轨迹达到tau2-bench的任一缺陷;在1120条为隔离电信缺陷而构建的路径上(该数量由构造决定),评估器奖励对挂起线路的重新加油,并使修复后的工具失败。最清晰的案例是一个临床基准测试,其工具告知智能体每次写入都在文档化的无写入设计下执行,而该设计其接口并未披露;其评分器将该消息作为证据,因此其行动成功率记录的是请求是否携带了预期载荷,而非是否有任何记录发生改变。

英文摘要

Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool call reports having done, assuming the tool did what its interface advertises. The audit taxonomies we survey publish no category for that assumption, and a defect beneath a score is present on every rerun. We treat a tool's advertised surfaces as an executable contract, check the implementation against it, and trace each score's provenance through the task files and evaluator code to the verdicts that derive from state a defective tool should have written. Across 34 audited mutating tools in four benchmarks we confirm seven tool defects and one evaluator property at pinned commits. On injected defects the checker raised no false positive in 25 flags, flagged 2 of 5 negative controls, and missed most: in 29 of 33 scored misses a clause covered the defect but no probe revealed it. The checker's own static half, run alone, flags 14 of 17 confirmed sites, so on these findings the dynamic half confirms and traces rather than discovers. Twelve further AgentDojo tools, with six held-out tools and the seven audited first, complete its 25-tool mutating surface, on which at least 5 tools diverge from their advertised surface as our contracts read it, a rate for AgentDojo alone. No gold trajectory reaches either tau2-bench defect; on 1,120 paths built to isolate the telecom defect, a number fixed by construction, the evaluator rewards a refuel of a suspended line and fails the repaired tool. The clearest case is a clinical benchmark whose tool tells the agent each write executed under a documented no-write design its interface does not disclose; its grader takes that message as evidence, so its action success rate records whether a request carried the expected payload, not whether any record changed.

补充信息

↑