arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29128cs.AIcs.LGcs.SE

APIFlow-Bench:评估智能体能否完成长时程、存在依赖关系的API工作流

APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows

Zelin Wan, Arash Nourian, Xiaoxiao Li, Nihar Nandan, Kamalakannan Nandagopal

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出APIFlow-Bench基准,评估智能体完成长时程依赖API工作流的能力,实验发现依赖链长度、模型可靠性影响成功率,复合故障不符合独立误差假设。

中文摘要 AI 辅助

使用工具的智能体通常通过一个二元指标评估:是否完成了端到端工作流。该指标无法区分生产环境中重要的故障,如凭证过期、有效载荷格式错误,或执行正确但最终交付错误的情况。我们推出APIFlow-Bench,这是一个可完全审计的基准,针对长时程、存在依赖关系的REST-API工作流,将性能分解为7项工程能力,并要求智能体生成由实际调用路径支持的答案。我们逐步生成合成API世界的子任务;每个子任务仅在零大语言模型(LLM)自测三元组验证其评分器、预言机确认可解性,且对抗性审计识别并修复6个评分器漏洞后才被接受。评分是确定性且依赖来源的:状态检查追踪模拟生成的金丝雀(canary)通过API数据流到答案必须源自的响应,且类型化答案卡片会被逐字段验证。我们发布所有答案密钥及44362份未编辑的执行记录。在19个前沿及开放权重模型(置于同一中立框架下)中,我们发现:(1)更长的依赖链会降低成功率,从单个子任务的93%降至干净的20子任务链的74%,若包含8%的链试验(经模型共识筛选后无模型判定为通过)则降至61%;(2)可靠性比最佳情况能力更能区分模型,五选一时跨度为7个百分点,但五选五的可靠性跨度达44个百分点;(3)复合故障的独立误差解释不符合数据:20子任务链的通过率比子任务级率的乘积高33个百分点,在干净切片中,77%的失败运行达到了正确的最终状态,仅在交付环节失败。

英文摘要

Tool-using agents are commonly evaluated by a single bit: whether an end-to-end workflow completed. This metric fails to distinguish failures that matter in production, such as expired credentials, malformed payloads, or correct execution followed by incorrect final delivery. We introduce APIFlow-Bench, a fully auditable benchmark for long-horizon, dependent REST-API workflows that decomposes performance into seven engineering capabilities and requires agents to produce answers supported by the actual call path. We generate synthetic API worlds forward, subtask by subtask; each subtask is admitted only after a zero-LLM self-test triad verifies its grader and an oracle establishes solvability, and an adversarial audit identified and fixed six grader exploits. Grading is deterministic and provenance-sensitive: a state check traces a mock-minted canary through the API data flow to the response the answer must originate from, and a typed answer card is verified field by field. We release all answer keys and 44,362 unredacted execution transcripts. Across 19 frontier and open-weight models under one neutral scaffold, we find: (1) longer dependency chains degrade success, from 93% on individual subtasks to 74% on clean 20-subtask chains and 61% when including the 8% of chain trials that a model-consensus screen flags as passed by no model; (2) reliability separates models more than best-case capability, with best-of-five spanning seven points but all-five-of-five reliability spanning 44 points; (3) the independent-error account of compounding failure does not fit the data: pass rates on 20-subtask chains are 33 percentage points above the product of subtask-level rates, and on the clean slice 77% of failing runs reached the correct final state and failed only at delivery.

发表机构

  • Postman, Inc.(波斯特曼公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑