arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SilentProbe:测量作为智能体工具的生产级API中的静默故障

SilentProbe: Measuring Silent Failure in Production APIs Used as Agent Tools

Zongrong Li, Shengkun Ye, Feiyou Guo, Zuoyou Dang

arXiv 2609.00035首次发表:更新:

发表机构

Texas A&M University; Monid, Inc.(德克萨斯农工大学; Monid公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SilentProbe研究发现生产级API的自然语言约束易引发智能体调用的静默故障,将词汇纳入模式可消除该故障,修复仅需一行模式代码。

AI 中文摘要

当大语言模型(LLM)智能体调用生产级API时,它无法区分“查询未匹配任何内容”与“服务器未理解查询”这两种情况,两者均返回带可解析主体的HTTP 200状态码,既无异常可捕获,也无字段可用于分支判断。本文研究了何种因素可预测这两种情况的发生,以及其对智能体的影响。研究人员审计了2501份独立发布的OpenAPI文档中的721320个参数,发现7.5%的参数声明了枚举值,15.2%的参数声明了任何可机器检查的约束,而40.1%的文档用自然语言描述了其模式未编码的至少一项约束。通过单一聚合层Monid(该层发布模式并为每次调用返回运行标识符),研究人员针对来自27家供应商的实时商业端点执行了219个源自模式的扰动测试,发现是约束形式而非供应商身份可预测API的诚实性:可机器检查的约束在111次测试中均返回诚实错误,而仅用自然语言描述的约束在61次测试中出现了44次静默失败(p=2e-13)。随后,8个系列的12种模型在常规任务上访问这些端点,对于仅在描述中举例说明的词汇,所有模型在88次尝试中均遗漏了88次;而完整写出的词汇的使用正确率为88%至91%。运行完整智能体循环时,模型在12%的案例中检测到了产生的静默故障,0%的案例中修复了故障,41%的案例中向用户断言了假阴性,12%的案例中编造了数据。将词汇纳入模式可消除故障,使失败案例从88/88降至0/89。该修复仅需一行模式代码,而非更好的模型。代码、模式、扰动集、智能体对话记录及每次调用的运行标识符已在该https URL发布。

英文摘要

An LLM agent calling a production API cannot distinguish a query that matched nothing from a query the server did not understand. Both return HTTP 200 with a parsable body, no exception to catch and no field to branch on. We ask what predicts which one occurred, and what it does to the agent. Auditing 721,320 parameters across 2,501 independently published OpenAPI documents, we find that 7.5% declare an enumeration and 15.2% declare any machine-checkable constraint at all, while 40.1% of documents state at least one constraint in prose that their schema does not encode. Executing 219 schema-derived perturbations against live commercial endpoints from 27 vendors, reached through a single aggregation layer (Monid) that publishes a schema and returns a run identifier for every call, we find that constraint form, not vendor identity, predicts honesty: machine-checkable constraints yielded an honest error in 111 of 111 cases, prose-only constraints failed silently in 44 of 61 (p = 2e-13). Twelve models across eight families then met these endpoints on ordinary tasks. A vocabulary that the description merely exemplifies was missed by every model on 88 of 88 attempts, while vocabularies written out in full were used correctly 88 to 91% of the time. Running the full agent loop, models detected the resulting silent failure in 12% of cases, repaired it in 0%, asserted a false negative to the user in 41%, and invented a figure in 12%. Promoting the vocabulary into the schema removes the failure, from 88 of 88 to 0 of 89. The fix is one line of schema rather than a better model. Code, schemas, perturbation sets, agent transcripts and per-call run identifiers are released at https://github.com/Jasper0122/silentprobe.

Comments12 pages, 9 figures. Code and data: https://github.com/Jasper0122/silentprobe

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑