AI 中文总结
针对流匹配假设训练数据完全可观测的局限,提出含缺失数据的流匹配方法,将缺失坐标视为潜在变量并平均流匹配损失,通过理论分析和实验验证其有效性,可应对实际数据缺失问题。
AI 中文摘要
流匹配假设训练数据完全可观测,但许多实际应用场景很少能满足这一条件。我们提出含缺失数据的流匹配方法,将训练样本的缺失坐标视为潜在变量,并对其可能取值上的流匹配损失进行平均。我们首先证明这种修正属于精确修正而非近似修正:在完全随机缺失且采用真实补全的情况下,不完整数据的目标函数与完整数据的目标函数完全等价,因此缺失性不会改变流匹配的学习内容,全部难点转移到补全模型上。随后,我们的有限样本分析解答了算法未明确的设计问题,且所得答案与直觉推断的结论不同:缺失性会转移估计量方差而非新增方差,每个样本仅需一次补全即可与完整数据的方差完全匹配,在固定评估预算下一次补全是最优选择;学习得到的补全模型会产生单一的不可约偏差,我们通过其与真实补全分布的条件Wasserstein距离期望对该偏差进行了界定。实验从数值上验证了理论预测,表明是确定性补全而非冻结补全会导致生成分布崩溃,且我们的方法在真实表格数据上可与强大的经典及深度补全基线方法相媲美。
英文摘要
Flow matching assumes fully observed training data, which many real-world applications rarely provide. We propose Missing-Data Flow Matching, which treats the missing coordinates of training samples as latent variables and averages the flow matching loss over the values they could take. We first prove the correction is exact rather than approximate. Under missing completely at random with true completions, the incomplete-data objective equals the complete-data objective, so missingness changes nothing about what flow matching learns and the entire difficulty relocates to the completion model. Our finite-sample analysis then answers design questions that the algorithm leaves open, and the answers are not the ones intuition suggests. Missingness transfers estimator variance rather than adding it, one completion per example already matches complete-data variance exactly, and under a fixed evaluation budget one completion is optimal. A learned completion model contributes a single irreducible bias, which we bound by its expected conditional Wasserstein distance to the true completion law. Experiments numerically validate the theoretical predictions, show that deterministic rather than frozen imputation is what collapses the generated distribution, and place our method alongside strong classical and deep imputation baselines on real tabular data.