Lookahead-R:通过执行中心规划进行预算感知的工具检索
Lookahead-R: Budget-Aware Tool Retrieval via Execution-Centric Planning
- Beijing University of Posts and Telecommunications(北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对工具检索中语义与执行验证的权衡,提出基于规划的Lookahead-R框架,利用执行感知世界模型和成本敏感MCTS,在ToolBench上实现91.40%的NDCG@5,优于现有方法。
AI中文摘要:
工具检索是基于大型语言模型的智能体在庞大且异构的API生态系统中运行时的关键瓶颈。现有方法面临固有的权衡:语义检索器速度快,但存在语义-功能差距;而基于执行的验证虽然提高了精确度,却以高昂的延迟为代价。我们提出Lookahead-R,一个基于规划的框架,将工具检索重新表述为资源受限的序列决策问题。其核心是,Lookahead-R引入了一个轻量级的执行感知代理世界模型,该模型在不调用真实API的情况下联合预测工具执行成功率、延迟成本和语义效用。该世界模型驱动一种成本敏感、不确定性引导的蒙特卡洛树搜索,在严格的预算约束下导航工具空间。在大规模ToolBench基准上的评估表明,Lookahead-R在所有测试场景中均实现了优越的准确率-效率权衡。在最具挑战性的I3分割上,它达到了91.40%的NDCG@5,比最先进的ToolGen(90.16%)高出1.24%。消融研究证实,显式延迟建模是在资源约束下识别高质量工具的关键判别信号。
英文摘要:
Tool retrieval is a critical bottleneck for LLM-based agents operating over large, heterogeneous API ecosystems. Existing approaches face an inherent trade-off: semantic retrievers are fast but suffer from the semantic-functional gap, while execution-based validation improves precision at the cost of prohibitive latency. We propose Lookahead-R, a planning-based framework that reformulates tool retrieval as a resource-constrained sequential decision-making problem. At its core, Lookahead-R introduces a lightweight execution-aware surrogate world model that jointly predicts tool execution success, latency cost, and semantic utility---without invoking real APIs. This world model drives a cost-sensitive, uncertainty-guided Monte Carlo Tree Search that navigates the tool space under strict budget constraints. Evaluated on the large-scale ToolBench benchmark, Lookahead-R achieves a superior accuracy-efficiency trade-off across all test scenarios. On the most challenging I3 split, it attains an NDCG@5 of 91.40\%, outperforming the state-of-the-art ToolGen (90.16\%) by 1.24\%. Ablation studies confirm that explicit latency modeling is the key discriminative signal for identifying high-quality tools under resource constraints.