arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38414cs.LG

无需训练的合成:面向表格、时序与关系型合成数据的纯推理流水线

Synthesis Without Training: An Inference-Only Pipeline for Tabular, Temporal, and Relational Synthetic Data

Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar

首次发表
浏览论文内容

中文总结 AI 辅助

提出GENSCRIPT纯推理流水线,通过统计画像和语言模型推断约束,无需训练即可生成表格、时序和关系型合成数据,兼顾效率与保真度。

中文摘要 AI 辅助

合成数据生成目前以“先拟合后采样”范式为主导:先在私有数据集上训练生成模型,再从中采样。尽管这一范式被广泛采用,但它面临三个挑战:(1)每个数据集都需要一次新的训练;(2)不同数据模态(如单表、时间序列和关系数据库)需要特定于任务的模型和特征工程;(3)生成的模型不透明,使其在数据约束下的行为难以检查。我们提出GENSCRIPT,一种消除模型训练的纯推理流水线。GENSCRIPT计算源数据的确定性统计画像(列类型、范围、缺失率、类别、相关性等),并将其(而非原始行)传递给语言模型以推断字段语义和跨列完整性约束。随后,一个编码智能体将画像和约束编译为可执行、可审计的采样器。这种统一方法无需特定于任务的建模即可支持单表、时序和关系数据。在四个单表基准上,GENSCRIPT在2分钟内构建生成器,并在6秒内采样5万行,同时在边际保真度方面与领先方法相差仅几个百分点。值得注意的是,它是唯一在Adult数据集中完美保留列间1对1映射的方法。在一个智能建筑数据集上,它生成的条件时间序列比两个基线更接近真实分布,并完美保留了相应关系数据库中的主键和外键关系。

英文摘要

Synthetic data generation is dominated by the fit-then-sample paradigm: a generative model is trained on a private dataset and then sampled from. Despite its widespread adoption, this paradigm faces three challenges: (1) a new training run is required for every dataset; (2) different data modalities, such as single tables, time series, and relational databases, require task-specific models and feature engineering; and (3) the resulting model is opaque, making its behavior under data constraints difficult to inspect. We propose GENSCRIPT, an inference-only pipeline that eliminates model training. GENSCRIPT computes a deterministic statistical profile of the source data (column types, ranges, missingness, categories, correlations, etc.) and passes it--rather than raw rows--to a language model to infer field semantics and cross-column integrity constraints. A coding agent then compiles the profile and constraints into an executable, auditable sampler. This unified approach supports single-table, temporal, and relational data without task-specific modeling. Across four single-table benchmarks, GENSCRIPT builds generators in 2 minutes and samples 50k rows within 6 seconds, while remaining within a few points of leading methods in marginal fidelity. Notably, it is the only method that perfectly preserves a 1-to-1 mapping between columns in the Adult dataset. On a smart-building dataset, it produces conditional time series that more closely match the real distribution than two baselines and perfectly preserves primary- and foreign-key relationships in the corresponding relational database.

发表机构

  • Betterdata AI
  • National University of Singapore(新加坡国立大学)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • KAJIMA Technical Research Institute Singapore(鹿岛技术研究所新加坡分所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑