arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14496cs.LGcs.AI

使用表格扩散Transformer生成基准健康数据

Generating Benchmark Health Data Using a Tabular Diffusion Transformer

Hao Yan, Lisa Pilgram, Dan Liu, Linglong Kong, Fida Dankar, Khaled El Emam

AI总结:

针对现有合成表格数据生成方法难以处理多异构表格的局限,本文提出两阶段跨表格数据生成框架,结合扩散Transformer实现高保真合成数据生成,验证了其有效性。

AI中文摘要:

跨表格数据生成(CTDG)旨在从多个异构表格中学习生成模型,并生成新的合成表格数据集。然而,现有的合成表格数据生成方法大多局限于单输入表格场景,难以有效处理具有不同特征集的多个异构表格。为解决这一局限,我们提出一种用于跨表格数据生成的两阶段框架。第一阶段,将每个异构原始表格转换为标准化统计表格,所有表格的列集相同,每个统计表格捕获原始列的边际分布及其两两相关性。第二阶段,训练扩散Transformer模型以捕获这些同质统计表格的结构模式,并生成合成统计表格。随后通过多元高斯采样及逆概率积分变换,从生成的统计表格中重构合成原始表格。该两阶段CTDG框架可从多个异构表格学习统一生成模型,并支持生成无限数量的真实合成异构表格。实验结果表明,所学习的统计表示具有高保真度,生成的合成数据在保真度-多样性权衡方面表现良好,验证了所提方法的有效性。

英文摘要:

Cross-Tabular Data Generation (CTDG) seeks to learn a generative model from multiple heterogeneous tables and produce new synthetic tabular datasets. However, existing synthetic tabular data generation methods are largely restricted to single-input-table scenarios and struggle to effectively handle multiple heterogeneous tables with diverse feature sets. To address this limitation, we propose a two-stage framework for cross-tabular data generation. In the first stage, each heterogeneous raw table is transformed into a standardized statistical table with the same set of columns across all tables. Each statistical table captures the marginal distributions of the original columns and the pairwise correlations among them. In the second stage, a diffusion transformer model is trained to capture structural patterns across these homogeneous statistical tables and to generate synthetic statistical tables. Synthetic raw tables are subsequently reconstructed from the generated statistical tables via multivariate Gaussian sampling followed by an inverse probability integral transform. This two-stage CTDG framework enables the learning of a unified generative model from multiple heterogeneous tables and supports the generation of an unlimited number of realistic synthetic heterogeneous tables. Experimental results demonstrate high fidelity in the learned statistical representations and a favorable fidelity-diversity trade-off in the generated synthetic data, validating the effectiveness of the proposed approach.

↑