arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Turnslide:通过遍历有限状态机实现可扩展的多轮数据合成

Turnslide: Scalable Multi-Turn Data Synthesis by Walking a Finite-State Machine

Aaron Fainman, Gabriela Kadlecová, Maciej Gryka, Bartosz Kruszczyński, Usman Zafar, Cédric Archambeau, Aaron Klein, David Salinas, Selim Nowicki, Jacek Golebiowski

arXiv 2610.07070首次发表:更新:

发表机构

Imperial College London; distil labs; Agon; ELLIS Institute Tübingen(伦敦帝国理工学院; distil 实验室; Agon 公司; ELLIS 图宾根研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于有限状态机的轻量级多轮数据合成框架,将API建模为状态机生成工具序列,一次LLM调用完成示例,微调SLM后准确率达70.7%,优于现有方法且token消耗减少3.6-6.6倍。

AI 中文摘要

小型语言模型服务成本低,且可在私有基础设施上运行,但基础模型在多轮工具调用方面往往不够出色,而对其进行微调需要每个API的专用数据,这类数据很少存在。现有的合成方法对于大规模微调而言过于昂贵,因为它们通常需要为不同领域模拟操作环境,并且每生成一轮对话就需要多次调用大语言模型。我们提出了一种全自动、轻量级的合成框架,将每个API建模为一个有限状态机,将系统表示为抽象状态,这些状态决定何时可以调用每个工具,从而生成状态有效的工具序列;这些序列通过一次大语言模型调用即可转化为完整的示例。我们不追求优化多样性,而是设定目标分布,覆盖轮数、工具序列和任务复杂度。我们通过在生成的轨迹上微调小型语言模型来衡量数据质量,结果表明,基于有限状态机的生成方法显著提升了下游准确率,相较于未变异的基线,在现有工作的对比中,达到了70.7%的完全准确率,而对比方法分别为63.4%和53.7%,且使用的token减少了3.6至6.6倍。

英文摘要

Small language models are inexpensive to serve and can run on private infrastructure, but base models are often not good enough at multi-turn tool calling, and fine-tuning them needs per-API data that rarely exists. Existing synthesis methods are too expensive for high-scale fine-tuning, as they often require mock operational environments for different domains and multiple LLM calls per generated conversation turn. We introduce a fully automated, lightweight synthesis framework that models each API as a finite-state machine, representing the system as abstract states that determine when each tool may be called, producing state-valid sequences of tools; sequences are translated into complete examples with a single LLM call. Rather than optimize diversity, we set a target distribution over the number of turns, the tool sequence and task complexity. We measure data quality by fine-tuning SLMs on generated trajectories, showing that our FSM-based generation significantly improves downstream accuracy over an unmutated baseline and, against existing works, reaches 70.7% full accuracy over 63.4% and 53.7% with 3.6-6.6$\times$ fewer tokens.

CommentsAccepted at the SLM-Agents Workshop, NeurIPS 2026 (non-archival)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑