arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过标签增强实现半监督条件扩散

Semi-Supervised Conditional Diffusion via Label Augmentation

Jin Su, Yuan Gao, Yong Zhou, Jian Huang

arXiv 2607.16685首次发表:更新:

发表机构

Department of Applied Mathematics, The Hong Kong Polytechnic University; School of Statistics and Data Science, Nankai University; School of Statistics, East China Normal University; Department of Data Science and AI, The Hong Kong Polytechnic University(应用数学系,香港理工大学; 统计与数据科学学院,南开大学; 统计学院,华东师范大学; 数据科学与人工智能系,香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何利用未标记数据改进条件扩散模型,提出标签增强条件扩散方法(LACD),通过给未标记数据赋平凡标签并联合去噪,在理论上保证可识别性和收敛速度,实验证明该方法在样本效率和生成性能上优于纯监督估计器。

AI 中文摘要

条件扩散模型已成为从标记数据中学习复杂条件分布的强大且灵活的框架。然而在实践中,获取高质量标签成本高且耗时,大量未标记数据未被利用。为解决此问题,我们引入标签增强条件扩散(LACD),一种简单有效的方法,通过为未标记示例分配指定的平凡标签并在增强数据集上进行联合去噪分数匹配来纳入它们。我们提供了在此方案下保证目标条件分布在总体水平上可识别的充分条件。此外,我们建立了严格的统计保证:当有足够多未标记样本时,LACD产生的采样分布在总变差距离上比纯监督估计器收敛得严格更快,在Wasserstein - 1距离上至少一样快。在合成、图像和表格基准上的大量实验证实了我们的理论,并显示出与纯监督估计器相比在样本效率和生成性能上有显著提升。

英文摘要

Conditional diffusion models have become a powerful and flexible framework for learning complex conditional distributions from labeled data. In practice, however, acquiring high-quality labels is costly and time-consuming, leaving large volumes of unlabeled data unused. To address this, we introduce label-augmented conditional diffusion (LACD), a simple and effective approach that incorporates unlabeled examples by assigning them a designated trivial label and performing joint denoising score matching over the augmented dataset. We provide sufficient conditions guaranteeing population-level identifiability of the target conditional distribution under this scheme. Moreover, we establish rigorous statistical guarantees: when sufficiently many unlabeled samples are available, the sampling distribution produced by LACD converges strictly faster than the purely supervised estimator in total variation distance, and at least as fast in Wasserstein-1 distance. Extensive experiments on synthetic, image, and tabular benchmarks corroborate our theory and show substantial gains in sample efficiency and generative performance compared with the purely supervised estimator.

Comments34 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑