面向混合类型数据的条件独立性正则化分布自动编码器
Conditional-Independence-Regularized Distributional Autoencoders for Mixed-Type Data
浏览论文内容
中文总结 AI 辅助
本研究提出Conditional-Independence-Regularized Distributional Autoencoders框架,通过结合不同类型变量的目标函数与条件独立性正则化项,实现混合类型数据的表征学习,在多类数据集上提升了分布恢复性能并保留了变量依赖结构。
中文摘要 AI 辅助
包含数值型与分类型变量的混合类型数据广泛存在于诸多科学及实际应用场景中。现有表征学习与生成建模方法通常聚焦于重建精度或无条件数据生成,却往往无法在恢复数据完整条件分布的同时,保留异质变量类型间可解释的结构关系。本研究提出Conditional-Independence-Regularized Distributional Autoencoders(条件独立性正则化分布自动编码器)这一框架,用于通过条件分布匹配与结构正则化学习混合类型数据的低维表征。该方法结合了针对数值型变量的基于能量得分的目标函数、针对分类型变量的基于似然的目标函数,以及辅助条件独立性正则化项,以鼓励学习到的表征捕捉数值型与分类型分量间的依赖关系。理论分析表明,最优表征需在未解释的数值变异性、分类型变量的条件熵与残差条件依赖之间取得平衡。实验结果显示,该方法在合成与真实数据集上均表现出色,大幅提升了分类型分布恢复效果,实现了具有竞争力的整体条件分布恢复,并保留了混合类型依赖结构,相关代码已在GitHub上公开。
英文摘要
Mixed-type data containing both numerical and categorical variables arise in many scientific and real-world applications. Existing representation learning and generative modeling approaches typically focus either on reconstruction accuracy or unconditional data generation, but often fail to recover the full conditional distribution of the data while preserving interpretable structural relationships between heterogeneous variable types. In this work, we introduce Conditional-Independence-Regularized Distributional Autoencoders, a framework for learning low-dimensional representations of mixed-type data through conditional distribution matching and structural regularization. Our method combines an energy-score-based objective for numerical variables, a likelihood-based objective for categorical variables, and an auxiliary conditional independence regularization term encouraging the learned representation to capture the dependence between numerical and categorical components. We provide theoretical analysis showing that the optimal representation balances unexplained numerical variability, conditional entropy of categorical variables, and residual conditional dependence. Empirically, the proposed method achieves strong performance on both synthetic and real-world datasets, substantially improving categorical distribution recovery, achieving competitive overall conditional distribution recovery, and preserving mixed-type dependence structure. The code has been made available at GitHub.
发表机构
- University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。