发表机构
Centre for the Science of Learning & Technology (SLATE), University of Bergen(卑尔根大学学习与技术科学中心(SLATE))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MIND通过边际传输和条件扩散分离边际与依赖建模,在混合类型表格生成中提升边际保真度和依赖保留,实现稳定平衡。
AI 中文摘要
本文提出MIND,一种用于混合类型表格数据的边际不变神经依赖扩散模型。MIND不直接在原始异构特征空间中学习联合分布。相反,它首先通过按列边际传输将不同的变量类型映射到统一的潜在依赖空间。然后,一个条件扩散模型学习跨列关系。Copula-切线去噪将已知的边际分量与可学习的依赖残差分离。采样阶段的秩投影进一步缓解反向扩散中的边际偏移。在九个不同的表格基准上的实验表明,MIND在边际保真度和依赖保留方面始终优于现有的统一方法。通过明确地将边际建模与依赖学习分离,MIND在边际保真度、联合依赖保留和下游预测效用之间实现了强大且稳定的平衡。这项工作支持将边际和依赖建模分离作为复杂混合类型表格生成的一种原则性且高效的方法。
英文摘要
This paper proposes MIND, a marginal-invariant neural dependency diffusion model for mixed-type tabular data. MIND does not directly learn the joint distribution in the original heterogeneous feature space. Instead, it first maps different variable types into a unified latent dependency space via column-wise marginal transport. A conditional diffusion model then learns cross-column relationships. Copula-tangent denoising separates known marginal components from learnable dependency residuals. Rank projection during the sampling phase further mitigates marginal shift in reverse diffusion. Experiments across nine diverse tabular benchmarks show that MIND consistently improves marginal fidelity and dependency preservation over existing unified approaches. By explicitly isolating marginal modelling from dependency learning, MIND achieves a strong and stable balance among marginal fidelity, joint dependency preservation, and downstream prediction utility. This work supports separating marginal and dependency modelling as a principled and highly effective paradigm for complex mixed-type tabular generation.