发表机构
University of Michigan, Ann Arbor; Microsoft Research, Redmond, WA(密歇根大学安娜堡分校; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AutoTailor通过离线过滤和在线动态重选,为Web智能体构建紧凑的MCP API集合,在保持或提升准确率的同时大幅降低令牌成本和延迟。
AI 中文摘要
Web智能体可以利用可复用的工具来降低底层浏览器交互的成本和延迟,但自动发现的工具集合可能规模庞大、冗余且与用户需求对齐不佳。我们提出AutoTailor,一个用于构建和维护紧凑的、从轨迹派生的模型上下文协议(MCP)API集合的元智能体框架。离线阶段,AutoTailor将Web轨迹转换为参数化的浏览器自动化程序,应用质量过滤器移除粒度不合适和功能冗余的API,并应用使用可能性过滤器优先保留广泛有用的能力,同时保持语义覆盖。在线阶段,动态重选机制监控任务结果和API使用情况,识别重复出现的覆盖缺口,添加相关候选,并修剪持续未使用的能力。我们在106个WebArena Postmill任务上评估AutoTailor。离线过滤将初始的1,283个未精炼API减少到87个,动态重选产生一个包含33个API的集合。结合推理与行动(ReAct)回退,该集合达到90.6%的正确率,而仅用ReAct为87.5%,同时平均总请求令牌成本降低57.8%,延迟降低29.4%。不使用ReAct时,它达到60.1%的正确率,与未精炼集合的性能基本持平,同时将请求令牌使用量减少94.9%。这些结果共同表明,静态过滤产生一个紧凑的API清单,预期支持核心、高可能性任务,而动态重选进一步根据观察到的用户需求定制该清单。这种组合提高了准确性和延迟,同时大幅减少令牌使用和端到端成本,展示了用户对齐能力管理对高效Web智能体的价值。
英文摘要
Web agents can utilize reusable tools to reduce the cost and latency of low-level browser interaction, but automatically discovered tool collections can be large, redundant, and poorly aligned with user demand. We present AutoTailor, a meta-agentic framework for constructing and maintaining a compact set of trajectory-derived Model Context Protocol (MCP) APIs. Offline, AutoTailor converts web trajectories into parameterized browser-automation programs, applies a Quality Filter to remove APIs with unsuitable granularity and redundant functionality, and applies a Usage Likelihood Filter to prioritize broadly useful capabilities while preserving semantic coverage. Online, Dynamic Reselection monitors task outcomes and API usage, identifies recurring coverage gaps, adds relevant candidates, and prunes persistently unused capabilities. We evaluate AutoTailor on 106 WebArena Postmill tasks. Offline filtering reduces the initial 1,283 unrefined APIs to 87, and Dynamic Reselection produces a 33-API set. With reasoning and acting (ReAct) fallback, this set achieves 90.6% correctness, compared with 87.5% for ReAct alone, while reducing average total request-token cost by 57.8% and latency by 29.4%. Without ReAct, it achieves 60.1% correctness, marginally matching the performance of unrefined set, while reducing request-token usage by 94.9%. Together, these results show that static filtering produces a compact inventory of APIs expected to support core, high-likelihood tasks, while dynamic reselection further tailors that inventory to observed user needs. This combination improves accuracy and latency while sharply reducing token usage and end-to-end cost, demonstrating the value of user-aligned capability management for efficient web agents.