ToolCompass:引导工具试用,而非抑制它
Toolcompass: Guiding Tool Trialing, Not Suppressing It
AI总结:
ToolCompass通过功能聚类引导工具试用,提升LLM智能体对未见工具的泛化,无需额外推理开销,在多个基准上显著提高任务成功率。
AI中文摘要:
大型语言模型(LLM)智能体必须将从训练期间见过的工具泛化到部署时未见过的工具。一个关键挑战是工具试用,即过度的试用会浪费交互预算,而选择性的试用则能探索不熟悉的工具。现有的基于结果的后训练让浪费的试用得不到引导,而回合级监督可能抑制必要的探索。我们引入了ToolCompass,一个后训练框架,通过根据共享功能组织工具调用表示来引导工具试用。具体来说,ToolCompass将每个功能类建模为von Mises--Fisher分布,并联合减少跨领域的类内变异,增加类间分离。这种结构将经验从见过的工具转移到功能相似的未见工具,引导探索远离不相关的替代方案。ToolCompass不需要真实调用轨迹或未见工具的访问权限,且不增加推理开销。在AppWorld和FTRL上的实验显示,在GRPO、RFT和DMPO上均有一致的提升。相比普通后训练,ToolCompass在AppWorld的OOD任务成功率上最高提升了10.71个百分点,并在两个基准上均优于竞争基线。
英文摘要:
Large language model (LLM) agents must generalize from tools seen during training to unseen tools at deployment. A key challenge is tool trialing, i.e., excessive trials waste the interaction budget, whereas selective trials enable exploration of unfamiliar tools. Existing outcome-based post-training leaves wasteful trials unguided, while turn-level supervision may suppress necessary exploration. We introduce ToolCompass, a post-training framework that guides tool trialing by organizing tool-call representations according to shared functions. Specifically, ToolCompass models each function class as a von Mises--Fisher distribution and jointly reduces intra-function variation across domains and increases inter-function separation. This structure transfers experience from seen tools to functionally similar unseen tools, directing exploration away from unrelated alternatives. ToolCompass requires no ground-truth call traces or unseen-tool access and incurs no inference overhead. Experiments on AppWorld and FTRL show consistent gains across GRPO, RFT, and DMPO. improves AppWorld OOD task success by up to 10.71 percentage points over vanilla post-training and performs best among competitive baselines on both benchmarks.