arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26836cs.AIcs.SE

智能体-工具交互中的静默失败:对ToolUniverse的审计

Silent Failures in Agent-Tool Interaction: An Audit of ToolUniverse

发表机构AI科技伦理
查看机构详情
  • AI Tech Ethics(AI科技伦理)

机构由 AI 辅助整理,请以论文原文为准。

Shreya Gopalan, Devansh Singh, Sundaraparipurnan Narayanan

首次发表
浏览论文内容

中文总结 AI 辅助

本研究审计ToolUniverse中15个科学工具的智能体-工具交互,发现91个静默失败(多源于API或包装器层),提出上下文可靠性概念以应对此类失败。

中文摘要 AI 辅助

智能体AI系统正日益采用集成多种工具的自动化流水线。虽然先前的研究和基准测试已关注这些智能体系统的任务成功率和任务完成情况,但关于智能体与工具交互的研究,特别是在生物学智能体工作流中的研究仍然有限。本研究调查了智能体与工具交互中的特定失败,即工具调用看似成功,但通过API/包装器从工具获取的部分或全部信息或功能不完整或缺失,且用户或智能体未收到关于此类缺失信息的任何通信或通知。我们将此称为静默失败,因为用户或智能体未意识到此类失败已发生。为进行本研究,我们开发了一种审计机制来识别智能体与工具交互中的此类静默失败,通过检查集成在ToolUniverse环境中的15个科学工具(及其相关API文档和工具文档)(ToolUniverse作为我们的实验环境而非研究对象本身)。我们围绕7个失败位点构建研究,这些位点表征失败在链条中发生的位置。我们观察到91个失败(在基于LLM的候选发现和自动化测试后进行人工验证),其中最常见的是数据或字段缺失以及搜索、过滤或排序标准中的不一致。91个失败中的大多数发生在API层(51个)或包装器层(25个),具有静默失败在下游放大的可能性。结果表明,静默失败起源于事件上游,并向下游传播为看似有效的科学输出。我们提出了一种上下文可靠性的概念来处理此类失败,并建议在智能体-工具交互流水线中测试、披露、监控和衡量此类失败的机制。

英文摘要

Agentic AI systems are increasingly adopting automated pipelines that integrate multiple tools. While prior research and benchmarks have studied about task success and task completion of these agentic systems, the research about agent to tool interaction, specifically in biology agentic workflow is limited. This study investigates specific failures in agent to tool interaction where a tool invocation appears successful, some or all of the information or functionality from the tool via API/ wrapper is incomplete or missing and there are no communications / notifications to the user or the agent about such missing information. We call this a silent failures as the user or the agents are not aware that such failure has occurred. For the purposes of this study we developed an audit mechanism to identify such silent failures in Agent to tool interaction, by examining 15 scientific tools (and their associated API documentation and tool documentations) integrated within ToolUniverse environment (ToolUniverse serves as our experimental environment rather than the object of the study itself). We structure our study around 7 failure locus characterising where the failure occurs in the chain. We observed 91 failures (manually validated post LLM based candidate discovery and automated testing), most frequent of them being missing data or fields and inconsistencies in search, filtering or ranking criteria. Most of the 91 failures occurred in API layer (51) or wrapper layer (25), with a potential of silent failure amplification downstream. The results show that silent failures originate upstream of the event and propagate downstream into apparently valid scientific outputs. We propose a concept of contextual reliability to handle such failures and suggest mechanisms for testing, disclosing, monitoring, and measuring such failures across the agent-tool interaction pipeline.

↑