菜单即执行先验:面向在线智能体的状态-路径工具菜单
The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents
- University of Central Florida(中佛罗里达大学)
- University of Rochester(罗切斯特大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对在线智能体面临海量工具接口的问题,提出状态-路径工具菜单,将菜单作为执行先验,通过编码器、检索器和重排序器选择并排序工具,在ToolBench上将成功率从0.737提升至0.898,且优于现有基线。
AI中文摘要:
语言模型通过工具行动,然而实际智能体面对的是包含数千个接口的库。我们引入工具菜单的概念,即执行前展示给智能体的可用工具的有序短子集。智能体只能调用此菜单中的工具。多步骤任务需要最终动作以及按可用顺序创建其输入的先决工具。当前的构造器按请求相关性对工具排序,这可能会突出最终动作,同时遗漏或延迟不太明显的生产者。我们引入状态路径,即从可观察的请求状态到期望结果的执行前路线,并提出状态-路径工具菜单来学习它。我们的框架将菜单视为这些路线上的执行先验。其编码器表示哪些工具可以从当前状态运行,它们的输出如何满足后续输入,以及哪些顺序在训练路径中重复出现。检索器覆盖可执行的入口、缺失输入的生产者和最终动作。然后,重排序器将生产者置于消费者之前。在ToolBench上,我们的菜单将在线成功率从0.737提升到0.898,并且在不改变智能体的情况下优于检索、重排序、生成和路由基线。状态-路径菜单用32个工具覆盖的完整链比官方列表用128个工具覆盖的更多,并且其成功增益在具有不同模型容量的执行器家族中持续存在。我们的代码位于此https URL。
英文摘要:
Language models act through tools, yet practical agents face libraries containing thousands of interfaces. We introduce the tool menu as the short, ordered subset of available tools shown to an agent before execution. The agent can call only tools in this menu. Multi-step tasks require the final action and the prerequisite tools that create its inputs in a usable order. Current constructors rank tools by request relevance, which can surface the final action while omitting or delaying less obvious producers. We introduce the state path, a pre-execution route from the observable request state to the desired outcome, and propose State-Path Tool Menu to learn it. Our framework treats the menu as an execution prior over these routes. Its encoder represents which tools can run from the current state, how their outputs satisfy later inputs, and which orders recur in training paths. A retriever covers an executable entry, the missing-input producers, and the final action. A reranker then places producers before consumers. On ToolBench, our menu raises online success from 0.737 to 0.898 and outperforms retrieval, reranking, generation, and routing baselines without changing the agent. The State-Path menu also covers more complete chains with 32 tools than the official list covers with 128, and its success gain persists across executor families with different model capacities. Our code is at https://github.com/Met2348/State-Path.