无需训练的合成:面向表格、时序与关系型合成数据的纯推理流水线
Synthesis Without Training: An Inference-Only Pipeline for Tabular, Temporal, and Relational Synthetic Data
浏览论文内容
中文总结 AI 辅助
提出GENSCRIPT纯推理流水线,通过统计画像和语言模型推断约束,无需训练即可生成表格、时序和关系型合成数据,兼顾效率与保真度。
中文摘要 AI 辅助
合成数据生成目前以“先拟合后采样”范式为主导:先在私有数据集上训练生成模型,再从中采样。尽管这一范式被广泛采用,但它面临三个挑战:(1)每个数据集都需要一次新的训练;(2)不同数据模态(如单表、时间序列和关系数据库)需要特定于任务的模型和特征工程;(3)生成的模型不透明,使其在数据约束下的行为难以检查。我们提出GENSCRIPT,一种消除模型训练的纯推理流水线。GENSCRIPT计算源数据的确定性统计画像(列类型、范围、缺失率、类别、相关性等),并将其(而非原始行)传递给语言模型以推断字段语义和跨列完整性约束。随后,一个编码智能体将画像和约束编译为可执行、可审计的采样器。这种统一方法无需特定于任务的建模即可支持单表、时序和关系数据。在四个单表基准上,GENSCRIPT在2分钟内构建生成器,并在6秒内采样5万行,同时在边际保真度方面与领先方法相差仅几个百分点。值得注意的是,它是唯一在Adult数据集中完美保留列间1对1映射的方法。在一个智能建筑数据集上,它生成的条件时间序列比两个基线更接近真实分布,并完美保留了相应关系数据库中的主键和外键关系。
英文摘要
Synthetic data generation is dominated by the fit-then-sample paradigm: a generative model is trained on a private dataset and then sampled from. Despite its widespread adoption, this paradigm faces three challenges: (1) a new training run is required for every dataset; (2) different data modalities, such as single tables, time series, and relational databases, require task-specific models and feature engineering; and (3) the resulting model is opaque, making its behavior under data constraints difficult to inspect. We propose GENSCRIPT, an inference-only pipeline that eliminates model training. GENSCRIPT computes a deterministic statistical profile of the source data (column types, ranges, missingness, categories, correlations, etc.) and passes it--rather than raw rows--to a language model to infer field semantics and cross-column integrity constraints. A coding agent then compiles the profile and constraints into an executable, auditable sampler. This unified approach supports single-table, temporal, and relational data without task-specific modeling. Across four single-table benchmarks, GENSCRIPT builds generators in 2 minutes and samples 50k rows within 6 seconds, while remaining within a few points of leading methods in marginal fidelity. Notably, it is the only method that perfectly preserves a 1-to-1 mapping between columns in the Adult dataset. On a smart-building dataset, it produces conditional time series that more closely match the real distribution than two baselines and perfectly preserves primary- and foreign-key relationships in the corresponding relational database.
发表机构
- Betterdata AI
- National University of Singapore(新加坡国立大学)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- KAJIMA Technical Research Institute Singapore(鹿岛技术研究所新加坡分所)
机构由 AI 辅助整理,请以论文原文为准。