arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29250cs.CL

Data Turnstile:面向函数调用数据生成的可扩展开源框架

Data Turnstile: A Scalable Open Framework for Function-Calling Data Generation

Goutham Ramakrishnan, Megha Sharma

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出开源框架Data Turnstile,可基于用户定义API规范生成高质量函数调用合成训练数据,经其微调的小型语言模型在BFCL、τ²-bench等基准上性能大幅提升,缩小了与更大规模模型的差距。

中文摘要 AI 辅助

小型语言模型(SLMs)因低延迟、低成本和端侧隐私特性,在智能体部署中颇具吸引力,但它们在工具使用任务中表现不佳,这类任务的训练数据稀缺且存在噪声。与更大规模的模型不同,SLMs无法凭借庞大的模型容量弥补低质量监督的不足,因此数据质量成为关键瓶颈。我们提出Data Turnstile,这是一个开源框架,可接收用户定义的API规范并生成用于函数调用的高质量合成训练数据。Turnstile将多轮工具使用交互分解为受约束的分步生成过程,同时引入验证与错误反馈循环,实现对API多样性、对话复杂度及输出正确性的细粒度控制。我们在两个具有挑战性的函数调用基准上验证了基于Turnstile生成数据的领域适配效果:在BFCL单轮基准上,使用Turnstile数据微调的Qwen3-0.6B模型,无需思维链即可达到75.9%的整体准确率,而启用思维链的基础模型准确率为67.4%;尽管该模型规模分别比启用思维链的Qwen3-1.7B和Qwen3-4B小3倍和7倍,但它缩小了与这两个模型的差距,后两者的准确率分别为78.4%和79.9%。在多轮智能体基准τ²-bench上,经Turnstile训练的Qwen3-1.7B模型在电信领域的通过率为31.1%,相比其基础模型的6.6%提升了4.7倍,且超过了规模大19倍的Qwen2.5-32B-Instruct模型(27.4%);经Turnstile训练的Qwen3-0.6B模型准确率为24.6%,相比其基础模型的3.5%提升了7倍,接近规模大53倍的32B模型。我们发布了Data Turnstile框架,以及包含1000+个API和10万+轮多轮交互的数据集。

英文摘要

Small language models (SLMs) are attractive for agentic deployment due to low latency, reduced cost, and on-device privacy, yet they struggle with tool-use tasks where training data is scarce and noisy. Unlike larger models, SLMs cannot compensate for low-quality supervision through sheer capacity, making data quality the critical bottleneck. We present Data Turnstile, an open-source framework that takes user-defined API specifications and generates high-quality synthetic training data for function calling. Turnstile decomposes multi-turn tool-use interactions into constrained, stepwise generation with validation and error-feedback loops, providing fine-grained control over API diversity, conversation complexity, and output correctness. We demonstrate effectiveness of domain adaptation with Turnstile data on two challenging function calling benchmarks. On the BFCL single-turn benchmark, a Qwen3-0.6B fine-tuned on Turnstile data without chain-of-thought achieves 75.9% overall accuracy (versus 67.4% for the base model with thinking enabled), closing the gap with thinking-enabled Qwen3-1.7B (78.4%) and Qwen3-4B (79.9%) despite being 3$\times$ and 7$\times$ smaller respectively. On $τ^2$-bench, a multi-turn agentic benchmark, Turnstile-trained Qwen3-1.7B achieves 31.1% pass^1 on the Telecom domain, improving 4.7$\times$ over its 6.6% base and surpassing Qwen2.5-32B-Instruct (27.4%), a model 19$\times$ larger. Turnstile-trained Qwen3-0.6B achieves 24.6%, improving 7$\times$ over its 3.5% base and approaching the 32B model (53$\times$ larger). We release Data Turnstile along with a dataset spanning 1,000+ APIs and 100K+ multi-turn interactions.

发表机构

  • Amazon AGI(亚马逊AGI)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑