arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ToolLIFT:将工具特定轨迹提升为函数级图以实现可泛化的工具规划

ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning

Xiuhui You, Jiayi Luo, Zichao Shen, Qingyun Sun, Ziwei Zhang

arXiv 2608.03468首次发表:更新:

AI 中文总结

本文提出ToolLIFT框架,通过将工具特定轨迹提升为函数级工作流图,实现了更具泛化性的工具规划,在多个基准上优于现有方法,对未见过的工具集泛化能力强。

AI 中文摘要

历史工具使用轨迹为大语言模型(LLM)智能体规划与协调工具使用提供了宝贵经验。现有方法直接从这些轨迹构建工具级图,但生成的图仍与特定工具绑定,难以在不同工具集间泛化。为解决这一挑战,研究发现尽管涉及的工具存在差异,但类似任务常共享通用的函数级工作流结构,该结构可作为工具规划中更具可迁移性的抽象。基于此洞见,本文提出ToolLIFT框架,其将工具特定轨迹提升为函数级工作流图(FWG)以实现可泛化的工具规划。具体而言,首先提出轨迹提升机制,该机制在FWG中编码工作流结构并在工具间共享协作经验;其次,基于FWG的全局结构,引入解耦的工作流规划与工具选择,使单个工具选择与整体工作流对齐;最后,为确保可靠的工具数据流,采用强化学习(RL)并提出源门控与技能特定的奖励函数,以在工具调用间维护可溯源的信息流。在两个分布内(ID)基准和三个分布外(OOD)基准上的实验表明,ToolLIFT始终优于当前最优基线,展现出对未见过的工具集的强泛化能力。

英文摘要

Historical tool-use trajectories provide valuable experience for large language model (LLM) agents to plan and coordinate tool usage. Existing approaches directly construct tool-level graphs from these trajectories, but the resulting graphs remain tied to specific tools and are hard to generalize across tool sets. To tackle this challenge, we find that despite differences in the tools involved, analogous tasks often share a common function-level workflow structure, which serves as a potentially more transferable abstraction for tool planning. Based on this insight, we propose ToolLIFT, a framework that lifts tool-specific trajectories into a function-level workflow graph (FWG) for generalizable tool planning. Specifically, we first propose a trajectory-lifting mechanism that encodes workflow structures in the FWG and shares collaboration experience across tools. Then, building on the global structure of the FWG, we introduce decoupled workflow planning and tool selection to align individual tool choices with the overall workflow. Lastly, to ensure reliable tool dataflow, we adopt Reinforcement Learning (RL) and propose source-gated and skill-specific rewards to maintain source-traceable information flow across tool calls. Experiments on two in-distribution (ID) and three out-of-distribution (OOD) benchmarks show that ToolLIFT consistently outperforms state-of-the-art baselines, demonstrating strong generalization to unseen tool sets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑