arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11098cs.SEcs.AIcs.CL

AgentCheck:用于基于MCP的大型语言模型智能体的重现-干预-缓解工作台

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP

Aritra Mazumder, Nusrat jahan Lia

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对工具使用型大语言模型智能体在工具出现问题时开发者难以处理故障的情况,提出AgentCheck开源工作台,通过特定运行和重放方式形成重现-干预-确认循环,能对故障模式进行评估验证,提升智能体故障处理能力。

中文摘要 AI 辅助

使用工具的大型语言模型智能体大多在假设所有工具都正常工作的情况下进行评估。当工具超时、返回过时值或其描述在部署中被篡改时,开发者需要一种可控的方式来重现故障、测试修复并在部署前确认修复有效。我们提出了AgentCheck,一个将MCP服务器转变为干预界面的开源网络工作台。AgentCheck让智能体针对其真实工具运行并记录每个工具响应,然后用故障注入器扰动的响应重新运行智能体。匹配的工具调用从缓存中重放,后续工具调用在智能体出现分歧后实时进行。这产生了一个重现-干预-确认循环。评分有两部分:确定性的通过/失败规则,以及用于解释性标签的大型语言模型判断,并通过人工注释进行验证。在五个智能体上,最佳的通过了105/120个场景,最弱的仅通过77个。故障通常是无声的,是对错误工具输出的自信使用而非崩溃。在最弱的智能体上,重试缓解措施将超时错误故障的成功率从低至30%提高到100%,而陈旧数据故障无论采取何种缓解措施仍保持在十分之三到四左右。AgentCheck使这些故障模式在部署前可重现、可比且可验证。

英文摘要

Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its description poisoned in deployment, the developer needs a controlled way to reproduce the failure, test a fix, and confirm the fix worked before deployment. We present AgentCheck, an open-source web workbench that turns an MCP server into an intervention surface. AgentCheck runs an agent against its real tools and records every tool response, then re-runs the agent with the response perturbed by a fault (12 types) injector. Matching tool calls are replayed from cache, and later tool calls go live after the agent diverges. This yields a reproduce-intervene-confirm loop: the developer toggles a mitigation, re-runs against the identical fault, and sees if the failure goes away. Scoring has two parts: deterministic pass/fail rules, plus an LLM judge for interpretive labels, validated against human annotations. Across five agents, the best passes 105/120 scenarios and the weakest only 77. The failures are usually silent, confident use of incorrect tool outputs rather than crashes. On the weakest agent, a retry mitigation raises success on timeout error faults from as few as 30% of cases to 100%, whereas stale-data faults remain near 3-4 of 10 regardless of the mitigation. AgentCheck makes these failure modes reproducible, comparable, and verifiable before deployment.

发表机构

  • University of Utah(犹他大学)
  • University of Dhaka(达卡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑