发表机构
School of Statistics, University of Minnesota(明尼苏达大学统计学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究混合连续-分类表格数据表示问题,提出Logit坐标框架,将分类变量编码与数值变量结合,形成Logit流匹配和扩散公式,经模拟和真实数据验证,该框架在多方面表现良好,能改进分布指标等。
AI 中文摘要
混合连续-分类数据给连续生成模型带来了表示问题。流匹配和高斯扩散在欧几里得空间中运行,而分类法则位于概率单纯形上且可能高度不平衡。我们研究了一个Logit坐标框架,它将分类变量编码为平滑的自然参数,并将其与变换后的数值变量相结合。这产生了Logit流匹配和Logit扩散的常见公式。我们引入了一种混合分布差异,将分类边际误差与条件连续瓦瑟斯坦误差分开,并推导了稳定性界限和不平衡感知非参数率,将向量场或漂移误差与解码后的混合分布误差联系起来。控制模拟表明,缩放后的Logit坐标改进或匹配了独热编码坐标,特别是在严重的稀有单元不平衡情况下。在四个真实数据基准和每个数据集十个分割上,Logit FM在三个数据集上改进了主要分布指标,在Churn2上相当;块条件Logit FM持续改进了扁平模型;Logit扩散总体上优于或匹配独热扩散。
英文摘要
Mixed continuous--categorical data pose a representation problem for continuous generative models. Flow Matching and Gaussian diffusion operate in Euclidean spaces, whereas categorical laws lie on probability simplices and may be highly imbalanced. We study a logit-coordinate framework that encodes categorical variables as smoothed natural parameters and combines them with transformed numerical variables. This yields common formulations of Logit Flow Matching and Logit Diffusion. We introduce a mixed-distribution discrepancy separating categorical marginal error from conditional continuous Wasserstein error, and derive stability bounds and imbalance-aware nonparametric rates linking vector-field or drift error to decoded mixed-distribution error. Controlled simulations show that scaled-logit coordinates improve or match one-hot coordinates, especially under severe rare-cell imbalance. Across four real-data benchmarks and ten splits per dataset, Logit FM improves the primary distributional metrics on three datasets and is comparable on Churn2; Block-Conditional Logit FM consistently improves the flat model; and Logit Diffusion generally improves over or matches One-Hot Diffusion.