发表机构
Appier AI Research; National Taiwan University(沛星人工智能研究院; 台湾大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SMITH框架联合训练工具创建与使用,4B Qwen3经训练在多任务上取得最优表现,其编写的工具还提升了其他模型的推理性能
AI 中文摘要
工具增强型语言模型受限于人类编写的API;现有工具创建系统通过在推理时提示冻结的大语言模型(LLM)来解决这一问题,但编写工具的模型与使用工具的模型相互分离,且没有信号表明其生成的模式(schema)是可被调用的模式。我们提出SMITH(基于模式的多任务迭代工具打磨,Schema-grounded Multi-task Iterative Tool Honing),这是一个在单一策略内联合训练工具创建与工具使用的强化学习框架。每次滚动(rollout)要么是构建任务(从少量示例编写工具),要么是使用任务(调用池化工具处理留存问题)。三个独立的奖励轴分别捕获模式、代码和结果失败,因此每种失败模式都会贡献各自的梯度。使用SMITH在13个程序推理任务上训练的4B Qwen3,搭配精确验证器,在留存任务上达到79.8的宏平均准确率,是所有评估方法中表现最佳的,且优于未训练的30B-A3B工具编写器。它还在TabMWP-Hard上达到40.4,在域外GQA上达到42.6(比同主干推理时基线的最佳表现高出7.6),且无需任何视觉或表格训练数据。我们的4B模型编写的工具还提升了LFM-2.5-350M和Qwen3-30B-A3B在相同推理任务下的性能。
英文摘要
Tool-augmented language models are bounded by the APIs humans bothered to write; existing tool-creation systems patch this by prompting a frozen LLM at inference time, leaving the model that writes a tool decoupled from the one that uses it, with no signal that the schemas it produces are schemas it can invoke. We propose SMITH (Schema-grounded Multi-task Iterative Tool Honing), a reinforcement learning framework that jointly trains tool creation and tool use inside a single policy. Each rollout is either a build task (write a tool from a few examples) or a use task (invoke a pooled tool on a held-out question). Three separate reward axes catch schema, code, and outcome failures independently, so each failure mode contributes its own gradient. A 4B Qwen3 trained with SMITH on 13 procedural reasoning tasks with exact verifiers reaches 79.8 macro-average accuracy on held-out tasks, the best across all evaluated methods and ahead of an untrained 30B-A3B tool-writer. It also reaches 40.4 on TabMWP-Hard and 42.6 on out-of-domain GQA (+7.6 over the best same-backbone inference-time baseline), without any visual or tabular training data. Tools written by our 4B models also lifted the performance of LFM-2.5-350M and Qwen3-30B-A3B under same reasoning tasks.