发表机构
Adaption Labs(Adaption Labs)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对零数据场景,提出基于提示的Invent-A-Dataset系统,从描述生成高质量多样化后训练数据,在五个前沿API上质量提升17%、多样性提升19%,显著改善下游模型性能。
AI 中文摘要
构建数据集仍然是AI开发中最手动、最脆弱的部分之一。在本技术报告中,我们聚焦于现实世界从业者面临的最极端但也最普遍的场景:零数据环境。在此环境中,从业者没有任何关于他们想要学习的能力的数据。我们引入了Invent-A-Dataset,这是一个基于提示的系统,能够从数据集描述生成逼真且大规模的后训练数据集。我们针对五个前沿模型API评估了Invent-A-Dataset,包括Anthropic、Google、OpenAI、DeepSeek和Zai。在八种任务类型和高达20K样本的数据集规模下,Invent-A-Dataset显著优于其他方法,同时实现了最高质量(相对提升17%)和最多样化的样本(相对提升19%)。其多样性优势随着训练数据集规模的扩大而增强(从200样本时的持平到20K样本时的相对提升37%)。这转化为可观的下游训练收益,使得后训练模型性能大幅提升。与其他生成器微调相比,Invent-A-Dataset微调在不同后训练模型架构中 consistently 排名更高。
英文摘要
Building datasets remains one of the most manual and brittle parts of AI development. In this technical report, we focus on the most extreme but also most prevalent setting real world practitioners face: a zero data regime. Here, practitioners don't have any data for the capability they want to learn. We introduce Invent-A-Dataset which is a prompt based system to go from dataset description to realistic and large scale post-training datasets. We evaluate Invent-A-Dataset against five frontier model APIs including Anthropic, Google, Open AI, DeepSeek, Zai. Across eight task types and dataset sizes up to 20K samples, Invent-A-Dataset significantly outperforms with both the highest quality (17% relative gains) while simultaneously producing the most diverse samples (19% relative gains). Its diversity advantage widens with scale of training dataset size (from parity at 200 samples to 37% relative gains at 20K samples). This translates into considerable downstream training gains, resulting in far more performant post-trained models. Invent-A-Dataset fine-tune consistently ranks higher compared to other generator fine-tunes across different post-trained model architectures.