arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Agent Seer:基于规范理解的场景合成

Agent Seer: Synthesizing Scenarios from Specification Understanding

Harish Karumuri, Mahesh Vemula, David Lopes Pegna

arXiv 2608.26133首次发表:更新:

发表机构

Apple(苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Agent Seer仅用单个MCP规范,无需示例、实时工具访问或领域调优,即可合成高质量测试场景,在7个不同领域的MCP规范上实现完整工具覆盖,为AI智能体评估提供了可扩展方案。

AI 中文摘要

评估使用外部工具的AI智能体需要能够体现从业者如何组合工具并在多轮对话中迭代的真实测试场景。手动构建此类场景需要深厚的领域专业知识,无法跨工具生态系统扩展,且生成的静态基准无法跟踪不断演变的API。我们发现,工具规范——函数名、自然语言描述和类型化参数模式——已编码了足够的语义信息,可在无需手动整理或实时工具执行的情况下合成真实的评估场景。Agent Seer正是利用这些潜在信息构建:仅需单个模型上下文协议(MCP)规范,无需示例、无需实时工具访问、无需特定领域调优。该流程丰富原始模式,生成带有合成工具输出的分级场景,并将其扩展为基于模拟数据的多轮对话,展现出强大的工具调用正确性和对话连贯性。通过将此流程应用于涵盖不同领域和工具套件规模的7个MCP规范,评估工具调用正确性和对话连贯性。该流程在所有领域均实现了强大质量,在中小型规范上实现了完整的工具覆盖。分析得出两个发现:参数模式复杂性是质量变化的最强关联因素——工具套件规模的影响较小且呈正交关系;参数值准确性是不完美场景中的主要失败模式,这一维度是粗粒度名称匹配指标无法察觉的。

英文摘要

Evaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain expertise, does not scale across tool ecosystems, and produces static benchmarks that cannot track evolving APIs. We observe that tool specifications -- function names, natural-language descriptions, and typed parameter schemas -- already encode sufficient semantic information to synthesize realistic evaluation scenarios without manual curation or live tool execution. Agent Seer builds off this latent information: from a single Model Context Protocol (MCP) specification, with no examples, no live tool access, and no domain-specific tuning. This pipeline enriches raw schemas, generates graded scenarios with synthetic tool outputs, and expands them into mock-data-grounded multi-turn dialogues that exhibit strong tool-calling correctness and conversational coherence. Evaluation quality is measured by applying this pipeline on seven MCP specifications spanning diverse domains and tool-suite sizes and measuring the tool-calling correctness and conversational coherence. The pipeline achieves strong quality across all domains, with complete tool coverage on small and medium specifications. Two findings emerge within this analysis: parameter schema complexity is the strongest correlate of quality variation -- tool-suite size plays a smaller, orthogonal role -- and argument value accuracy is the dominant failure mode among imperfect scenarios, a sub-dimension invisible to coarse-grained name-match metrics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑