发表机构
CEREMADE CNRS; Université Paris Dauphine - PSL; Cornell University(CEREMADE 法国国家科学研究中心; 巴黎第九大学 - PSL; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探讨扩散模型在固定预算训练下,用合成数据替换真实数据导致模型坍缩,但保留部分特征,而完全替换则快速退化,揭示了两种协议下特征保留的差异。
AI 中文摘要
模型坍缩是指当生成模型在由先前模型产生的合成数据上训练时出现的现象。由于其在社会和技术层面的影响,该现象已引起广泛关注。然而,以往研究得出了看似矛盾的结论:用合成数据替换真实数据会导致坍缩(Shumailov等人),而将真实数据与合成数据累积在一起则可以防止坍缩。对于扩散模型,我们研究了有限预算流水线中典型的中间机制:保留所有过去的数据集和真实数据,但每个新模型从这一不断增长的池中训练固定大小的样本,因此真实数据的比例逐渐消失,而无需移除任何数据。在二维螺旋数据集以及图像基准(MNIST、Fashion-MNIST和CIFAR-10)上的实验表明,替换协议会如文献所述迅速降低数据集质量,而固定预算协议仅部分降低,保留了一些特征。通过随机递推分析的多代参数动力学的线性响应模型证实了两种协议之间的差异:在两种协议下,某些特征在几代内就会变得脆弱并丢失,而在固定预算协议下,某些特征则具有鲁棒性,并能在几乎无限的时间范围内得以保留。
英文摘要
Model collapse arises when generative models are trained on synthetic data produced by earlier models. The phenomenon has attracted considerable attention because of its societal and technical implications. However, previous studies have reached seemingly contradictory conclusions: replacing real data with synthetic data causes collapse (Shumailov et al.), yet accumulating real data alongside synthetic data can prevent it. For diffusion models, we study an intermediate regime typical of finite-budget pipelines: all past datasets and the real data are kept, but each new model is trained on a fixed-size sample from this growing pool, so the real fraction vanishes without any data being removed. Experiments on a 2D spiral dataset as well as the image benchmarks (MNIST, Fashion-MNIST, and CIFAR-10) show that replacement protocol degrades dataset rapidly as in the literature, whereas the fixed budget degrades only partially, sparing some features. A linear-response model of the multi-generational parameter dynamics, analyzed by stochastic recursion, confirms that the two protocols differ: some features will be fragile and lost within a few generations for both protocols, while some will be robust and preserved over practically unbounded horizons under the fixed budget protocol.
CommentsNeurips 2026 PriGM workshop paper