发表机构
DIRO, Université de Montréal; Mila - Quebec Artificial Intelligence Institute; Air Liquide(蒙特利尔大学DIRO学院; 米拉-魁北克人工智能研究所; 液化空气集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出单智能体强化学习模型,在ChemToolBench多工具化学任务上,相较CheMatAgent的树搜索方法,在Qwen-2.5-7B和Llama-3.1-8B上提升了工具F1、返回F1及答案通过率,且成本更低。
AI 中文摘要
化学问题常需要语言模型无法从参数中获取的精确计算和数据库查询,因此必须借助外部工具。工具使用是一个三部分问题:从大量工具中选择合适的工具,用正确类型的参数填充工具,以及链式调用使每个调用消耗上一个调用的输出。CheMatAgent是此前发布的系统,它采用分层进化MCTS,分离策略和执行模型,在两个学习到的评判器下搜索工具调用树,其中一个评判器部分回归到GPT分配的分数。我们证明单个策略就足够。我们的模型将推理、工具调用和返回交织在一个从左到右的生成过程中,通过监督预热训练,然后在直接从正确调用链读取的程序化奖励上进行结果级强化学习训练,这使得训练循环中没有学习到的评判器和评判者。在ChemToolBench多工具综合化学任务上,对于CheMatAgent使用的两个主干模型,与最强的搜索配置相比,我们在Qwen-2.5-7B上将工具F1提高了5.5%,返回F1提高了9.6%;在Llama-3.1-8B上将工具F1提高了3.7%,返回F1提高了3.9%,且每个问题仅需一次模型调用,而搜索的成本随树的规模增长;我们还在Qwen-2.5-7B上领先答案通过率。
英文摘要
Chemistry questions often demand exact computation and database lookups that a language model cannot supply from its parameters, so it must reach for external tools. Tool use here is a three-part problem: select the right tool from a large pool, fill it with correctly typed arguments, and chain calls so that each consumes the outputs of the last. CheMatAgent, a previously published system, addresses this with hierarchical evolutionary MCTS: separate policy and execution models searching tool-call trees under two learned critics, one regressed partly onto GPT-assigned scores. We show that a single policy suffices. Our model interleaves reasoning, tool calls, and returns in one left-to-right generation, trained by a supervised warm-up and then outcome-level reinforcement learning against a programmatic reward read directly off the gold call chain, which leaves no learned critic and no judge in the training loop. On ChemToolBench multiple-tool comprehensive chemistry, on both backbones CheMatAgent use, we improve Tool F1 by 5.5% and Return F1 by 9.6% on Qwen-2.5-7B, and by 3.7% and 3.9% on Llama-3.1-8B, compared with their strongest search configuration, at one model invocation per question, against a search whose cost grows with the tree; we also lead answer Pass Rate on Qwen-2.5-7B.