发表机构
Imperial College London; distil labs; Agon; ELLIS Institute Tübingen(伦敦帝国理工学院; distil 实验室; Agon 公司; ELLIS 图宾根研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于有限状态机的轻量级多轮数据合成框架,将API建模为状态机生成工具序列,一次LLM调用完成示例,微调SLM后准确率达70.7%,优于现有方法且token消耗减少3.6-6.6倍。
AI 中文摘要
小型语言模型服务成本低,且可在私有基础设施上运行,但基础模型在多轮工具调用方面往往不够出色,而对其进行微调需要每个API的专用数据,这类数据很少存在。现有的合成方法对于大规模微调而言过于昂贵,因为它们通常需要为不同领域模拟操作环境,并且每生成一轮对话就需要多次调用大语言模型。我们提出了一种全自动、轻量级的合成框架,将每个API建模为一个有限状态机,将系统表示为抽象状态,这些状态决定何时可以调用每个工具,从而生成状态有效的工具序列;这些序列通过一次大语言模型调用即可转化为完整的示例。我们不追求优化多样性,而是设定目标分布,覆盖轮数、工具序列和任务复杂度。我们通过在生成的轨迹上微调小型语言模型来衡量数据质量,结果表明,基于有限状态机的生成方法显著提升了下游准确率,相较于未变异的基线,在现有工作的对比中,达到了70.7%的完全准确率,而对比方法分别为63.4%和53.7%,且使用的token减少了3.6至6.6倍。
英文摘要
Small language models are inexpensive to serve and can run on private infrastructure, but base models are often not good enough at multi-turn tool calling, and fine-tuning them needs per-API data that rarely exists. Existing synthesis methods are too expensive for high-scale fine-tuning, as they often require mock operational environments for different domains and multiple LLM calls per generated conversation turn. We introduce a fully automated, lightweight synthesis framework that models each API as a finite-state machine, representing the system as abstract states that determine when each tool may be called, producing state-valid sequences of tools; sequences are translated into complete examples with a single LLM call. Rather than optimize diversity, we set a target distribution over the number of turns, the tool sequence and task complexity. We measure data quality by fine-tuning SLMs on generated trajectories, showing that our FSM-based generation significantly improves downstream accuracy over an unmutated baseline and, against existing works, reaches 70.7% full accuracy over 63.4% and 53.7% with 3.6-6.6$\times$ fewer tokens.
CommentsAccepted at the SLM-Agents Workshop, NeurIPS 2026 (non-archival)