arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13553stat.MEmath.PR

条件分布框架用于验证合成多元数据

A Conditional-Distribution Framework for Validating Synthetic Multivariate Data

Hari Dahal, Ishanu Chattopadhyay

中文总结 AI 辅助

提出基于全条件分布的模型无关框架,通过MAP对齐统计量和最近真实相似性诊断,区分合成数据中的依赖丧失、记录重用和模式集中三种失败模式,并在多个真实数据集上验证其有效性。

中文摘要 AI 辅助

合成多元数据的统计验证需要评估生成器是否保留了目标总体的联合依赖结构,而不仅仅是重现观察到的记录。我们开发了一个基于全条件分布的模型无关框架。对于每个坐标,我们将观察值所分配的条件概率除以同一记录上下文中可用的最大条件概率;平均该量得到一个单侧MAP对齐统计量,该统计量可以使用在保留的真实数据上拟合的条件模型进行估计。数学贡献有两方面:在严格正性和兼容性下,完整的归一化条件剖面识别联合分布,其积分L1差异定义了有限状态生成过程上的度量;我们还建立了相应经验估计量的一致性和有限样本集中性。由于仅高条件对齐可能源于复制或对条件模式的集中,我们将其与最近真实相似性配对,作为单独的记录级新颖性诊断。我们在NSHAP健康和老龄化数据、乙型流感基因组监测以及34个综合社会调查波次上评估了该框架。在GSS中,大型科学模型在平均条件对齐方面与原始数据对照组匹配,同时保持显著的新颖性,表明保留了条件结构而没有行重用。在乙型流感中,Chow-Liu生成器匹配了对照对齐但几乎没有新颖性,揭示了观察记录的近乎复制。因此,该框架区分了三种统计上不同的失败模式:依赖丧失、记录重用和模式集中,并为验证合成健康、监测和人口数据提供了原则性基础。

英文摘要

Statistical validation of synthetic multivariate data requires assessing whether a generator preserves the joint dependence structure of the target population without merely reproducing observed records. We develop a model-agnostic framework based on full conditional distributions. For each coordinate, we normalize the conditional probability assigned to the observed value by the largest conditional probability available in the same record context; averaging this quantity yields a one-sided MAP-alignment statistic that can be estimated using a conditional model fitted on held-out real data. The mathematical contribution is twofold: under strict positivity and compatibility, the complete normalized conditional profile identifies the joint distribution, and its integrated L1 difference defines a metric on finite-state generative processes; we also establish consistency and finite-sample concentration for the corresponding empirical estimators. Because high conditional alignment alone can arise from copying or concentration on conditional modes, we pair it with nearest-real similarity as a separate record-level novelty diagnostic. We evaluate the framework on NSHAP health and aging data, influenza B genomic surveillance, and 34 General Social Survey waves. In GSS, the Large Science Model matched the original-data control in mean conditional alignment while retaining substantial novelty, indicating preservation of conditional structure without row reuse. In influenza B, a Chow-Liu generator matched the control alignment but had almost no novelty, revealing near-reproduction of observed records. The framework therefore distinguishes three statistically different failure modes: loss of dependence, record reuse, and mode concentration, and provides a principled basis for validating synthetic health, surveillance, and population data.

↑