AI 中文总结
该研究指出因子化生成模型中仅匹配潜在风格的边际分布无法保证其与类别信息独立,通过实验证实存在风格泄漏,并提出相关缓解策略及验证方法。
AI 中文摘要
因子化生成模型通常通过将潜在风格变量z_s的边际分布与固定高斯先验匹配来对其进行正则化,并将此视为风格表示与类别信息独立的证据。我们表明这种解释是不正确的:仅匹配边际分布不会对类别条件分布施加任何约束,尽管整体上风格表现为完美的高斯分布,但潜在风格仍可高度预测标签。我们推导了一个精确分解,表明这种不匹配是因子化采样所需的四个条件之一,并证明消除它是获得预期因子化的必要但非充分条件。实验中,我们的案例研究模型和四个代表性潜在基线实现了接近零的全局MMD,同时仍允许线性探针以74%--100%的准确率恢复类别标签(随机概率为10%)。我们的模型达到99.15%的聚类准确率,而外部评估的类别条件生成仅成功16%的时间。这种泄漏在涉及模型容量、课程、先验几何和跨两个数据集监督的六种独立扰动下仍然存在。四种缓解策略将探针准确率降低至21%--46%,尽管它们在很大程度上保留了类内依赖关系。事后条件先验在不重新训练的情况下将MNIST上的外部评估类别生成提高至0.97,但在CIFAR-10上仅达到0.41,而经验风格库在CIFAR-10上达到0.88。这些结果表明,仅基于风格潜在变量边际分布计算的任何散度都无法证明其与类别标签的独立性,且仅报告边际统计量无法验证因子化生成模型中通常声称的属性。
英文摘要
Factorized generative models commonly regularize a latent style variable z_s by matching its marginal distribution to a fixed Gaussian prior and interpret this as evidence that the style representation is independent of class information. We show that this interpretation is incorrect. Matching only the marginal distribution places no constraint on the class-conditional distributions, allowing the latent style to remain highly predictive of the label despite appearing perfectly Gaussian in aggregate. We derive an exact decomposition showing that this mismatch is one of four conditions required for factorized sampling, and demonstrate that eliminating it is necessary but not sufficient to obtain the intended factorization. Empirically, our case-study model and four representative latent baselines achieve near-zero global MMD while still allowing a linear probe to recover class labels with 74%--100% accuracy (10% chance level). Our model reaches 99.15% clustering accuracy, whereas externally evaluated class-conditional generation succeeds only 16% of the time. This leakage remains under six independent perturbations involving model capacity, curriculum, prior geometry, and supervision across two datasets. Four mitigation strategies reduce probe accuracy to 21%--46%, although they leave within-class dependence largely unchanged. A post-hoc conditional prior improves externally evaluated class generation to 0.97 on MNIST without retraining but reaches only 0.41 on CIFAR-10, while an empirical style bank achieves 0.88 on CIFAR-10. These results demonstrate that no divergence computed solely on the marginal distribution of the style latent can certify independence from class labels, and that reporting marginal statistics alone does not verify the property commonly claimed in factorized generative models.