arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09572stat.MLcs.AIcs.LG

通过SGD在高维线性回归中利用合成数据进行学习

Learning with Synthetic Data via SGD in High-Dimensional Linear Regression

Jichu li, Difan Zou

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过高维线性回归中的SGD分析,发现混合训练合成数据导致模型崩溃,而两阶段训练可避免此问题,并提供了风险界、缩放定律及两阶段训练优于纯真实训练的充要条件。

中文摘要 AI 辅助

合成数据已成为将模型训练扩展到有限人类生成数据之外的一种有前景的方式,但它也可能导致严重的模型崩溃(Dohmatob等人,2024),即任何固定比例的合成数据都会阻止模型性能在数据扩展下提升,留下一个非消失的额外风险下限。在本文中,我们研究了在具有模型偏移的高维线性回归中,合成数据如何影响单遍SGD的泛化性能。我们为混合训练和两阶段训练建立了有限样本风险界,将标准偏差和方差与源不匹配效应(即混合训练下的波动和持续漂移,以及两阶段训练下的过滤初始化偏差)区分开来。这些界揭示了一个鲜明的对比:混合训练会导致严重的模型崩溃,而两阶段训练通过仅在第一阶段使用合成数据来避免风险下限,表明在简单的数据课程下崩溃并非不可避免。在随机草图模型下,我们进一步获得了两种协议的缩放定律,其中混合训练在优化饱和状态下具有紧的结果。这些定律表明,在混合训练下,更大的模型可能会放大合成数据引起的退化,并量化了高质量合成预训练如何在两阶段训练中减少偏差。最后,我们建立了在相同真实数据预算和相同真实阶段更新下,两阶段训练严格优于仅真实训练的精确有限样本充要条件。总体而言,我们的结果强调合成数据既非固有有害也非固有有益;其效果关键取决于其质量以及用于整合它的训练协议。

英文摘要

Synthetic data has become a promising way to scale model training beyond limited human-generated data but it may also induce strong model collapse (Dohmatob et al., 2024), where any fixed fraction of synthetic data prevents model performance from improving under data scaling, leaving a non-vanishing excess risk floor. In this paper, we study how synthetic data affects the generalization of one-pass SGD in high-dimensional linear regression with model shift. We establish finite-sample risk bounds for mixed and two-stage training, separating standard bias and variance from source-mismatch effects, namely fluctuation and persistent drift under mixing and filtered initialization bias under two-stage. These bounds reveal a sharp contrast: mixed training induces strong model collapse, while two-stage training avoids the floor by using synthetic data only in the first stage, showing that collapse is not inevitable under a simple data curriculum. Under a random sketch model, we further obtain scaling laws for both protocols, with tight results for mixed training in the optimization-saturated regime. These laws show that larger models may amplify synthetic-induced degradation under mixing, and quantify how high-quality synthetic pretraining may reduce bias in two-stage training. Finally, we establish an exact finite-sample necessary-and-sufficient condition for two-stage training to strictly outperform real-only training under the same real-data budget and identical real-stage updates. Overall, our results highlight that synthetic data is neither inherently harmful nor beneficial; its effect depends critically on both its quality and the training protocol used to incorporate it.

发表机构

  • Renmin University of China(中国人民大学)
  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

↑