AI 中文总结
本文针对生成式AI领域中模型坍缩(MC)现象的研究空白,综述了不同应用场景下MC的进展及应对措施,指出了相关挑战与未来研究机遇。
AI 中文摘要
在海量网络级数据的驱动下,生成式AI(GenAI)已取得显著进展,在多个领域实现了各类应用。生成式AI的进步促使从业者使用AI合成数据训练下一代AI模型。不可否认,使用合成数据缓解了日益严苛的数据供给需求,但也引入了一个新的关键问题:在模型与数据的自消耗循环中,模型最终会发生坍缩(model collapse,MC),引发生成式AI更多的可信性担忧。近年来,越来越多的研究调查了模型坍缩(MC)现象并探索了缓解该问题的潜在解决方案。然而,针对MC现象的综述仍处于空白状态。为填补这一缺口,本文对相关研究提供了最新概述,梳理并综述了不同应用场景下MC的研究进展及缓解MC的应对措施,同时指出了面临的挑战与未来研究机遇。
英文摘要
Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors. The advances of GenAI have actuated practitioners to use AI-synthesized data for training next-generation AI models. Undeniably, using synthetic data has alleviated the increasing stringent demand for data supply. Unfortunately, it also introduces a new critical issue: in a self-consuming cycle between model and data, the model ultimately collapse, raising more trustworthiness concerns to GenAI. In recent years, increasingly more studies have investigated the phenomenon of model collapse (MC) and explored potential solutions to mitigate it. However, the review of the phenomenon of MC still remains blank. To fill this gap, this paper provides an up-to-date overview of these studies for consolidating and reviewing the progress of MC in different application scenarios and countermeasures for mitigating MC. We also highlight challenges and future research opportunities.
Comments11 pages, 1 figure, Accepted and published in Proceedings of IEEE AAIML 2026
Journal refProceedings of the IEEE International Conference on Advances in Artificial Intelligence and Machine Learning, 2026
DOI:10.1109/AAIML67890.2026.11498213