发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出PRODiGI,一种预训练的数据到程序模型,可在单次前向传播中推断显式生成程序,支持直接采样和密度评估,并通过程序空间微调将生成误差降低84%,实现快速、可解释的表格生成建模。
AI 中文摘要
从有限样本集估计概率密度通常需要针对特定数据集进行模型拟合。我们提出了PRODiGI,一种预训练的数据到程序模型,它能在单次前向传播中推断出一个显式、可执行的生成式程序。通过在合成数据集及其对应的真实程序上进行预训练,PRODiGI通过模板预测和非自回归程序参数解码,适应了多样的生成族和数据维度。其推断出的程序支持直接采样、密度和分数评估,并且可以独立于预训练模型进行检查。我们进一步引入了程序空间微调,该方法通过匹配生成样本和经验样本,在不改变模型参数的情况下,优化可微分的程序参数。实验表明,与现有预训练模型相比,PRODiGI实现了更低的平均密度和分数平均绝对误差,同时比其最接近的竞争对手提供了多倍的加速。程序空间微调进一步将生成最大均值差异降低了84%。通过将经验数据转化为显式、可复用的程序,PRODiGI为快速、可解释的表格生成建模开辟了新方向。
英文摘要
Estimating probability densities from a finite set of samples typically requires dataset-specific model fitting. We introduce PRODiGI, a pretrained data-to-program model that infers an explicit, executable generative program in a single forward pass. Pretrained on synthetic datasets paired with their ground-truth programs, PRODiGI accommodates diverse generative families and data dimensionalities through template prediction and non-autoregressive program parameter decoding. Its inferred programs support direct sampling, density and score evaluation, and inspection independently of the pretrained model. We further introduce program-space fine-tuning, which refines differentiable program parameters by matching generated and empirical samples while keeping model parameters intact. Experiments show that PRODiGI achieves lower average density and score MAE than existing pretrained models, while offering multi-fold speedups over its closest competitors. Program-space fine-tuning further reduces generation MMD by 84%. By turning empirical data into explicit, reusable programs, PRODiGI introduces a new direction for fast, interpretable tabular generative modeling.
Comments51 pages, 20 figures,