CDMD:用于表格数据的跨数据集混合类型扩散模型
CDMD: A Cross-Dataset Mixed-Type Diffusion Model for Tabular Data
浏览论文内容
中文总结 AI 辅助
本文提出CDMD,一种跨数据集混合类型扩散模型,直接在混合特征空间上联合训练,通过模式受限反向过程与共享Transformer去噪器,在七个数据集上实现最优平均生成质量,并验证了预训练的可迁移性。
中文摘要 AI 辅助
表格数据的生成模型通常针对每个数据集单独训练,限制了知识迁移,并且需要存储许多专门的模型。在本文中,我们引入了CDMD,一种在具有不同模式(schema)以及不同数值和类别特征数量的异构数据集上联合训练的表格扩散模型。与现有在连续表示空间中操作的跨数据集表格扩散模型不同,CDMD直接在混合类型特征空间上定义扩散,并进行端到端训练。为了适应异构类别域,我们为掩码扩散模型引入了一种模式受限的反向过程参数化,其中输出空间动态适应每个特征的词汇表。然后,我们将数值和类别特征级扩散过程组合成依赖于模式的行级过程。一个共享的、模式感知的Transformer去噪器捕获特征之间的依赖关系,并在不同模式间参数化反向过程。在七个真实世界数据集上,一个单独联合训练的CDMD在强单数据集和跨数据集基线中实现了最高的平均生成质量,同时使用的总参数远少于单独训练的模型集合。此外,在337个数据集的语料库上进行预训练,在有限的目标数据和有限的适应轮次下,提高了在以前未见过的数据集上的生成质量。这些结果证明了直接混合类型扩散在共享和可迁移的表格数据生成方面的潜力。我们的代码可在以下网址获取:https://this.https.url。
英文摘要
Generative models for tabular data are typically trained separately for each dataset, limiting knowledge transfer and requiring the storage of many specialized models. In this paper, we introduce CDMD, a tabular diffusion model trained jointly across heterogeneous datasets with different schemas and variable numbers of numerical and categorical features. Unlike existing cross-dataset tabular diffusion models that operate in continuous representation spaces, CDMD defines diffusion directly over the mixed-type feature space and is trained end-to-end. To accommodate heterogeneous categorical domains, we introduce a schema-restricted reverse-process parameterization for masked diffusion models, in which the output space dynamically adapts to each feature's vocabulary. We then compose numerical and categorical feature-level diffusion processes into a schema-dependent row-level process. A shared schema-aware Transformer denoiser captures dependencies between features and parameterizes the reverse process across varying schemas. On seven real-world datasets, a single jointly trained CDMD achieves the highest average generation quality among strong single-dataset and cross-dataset baselines, while using substantially fewer total parameters than the collection of separately trained models. Furthermore, pre-training on a corpus of 337 datasets improves generation on previously unseen datasets under both limited target data and limited adaptation epochs. These results demonstrate the potential of direct mixed-type diffusion for shared and transferable tabular data generation. Our code is available at https://github.com/ketatam/cdmd.
发表机构
- Munich Data Science Institute, Technical University of Munich(慕尼黑数据科学研究所,慕尼黑工业大学)
- SAP SE(SAP公司)
机构由 AI 辅助整理,请以论文原文为准。