防止模型崩溃:从Fisher-Rao视角看合成数据训练的动态
Preventing Model Collapse: A Fisher-Rao Perspective on the Dynamics of Training with Synthetic Data
- University of California at Los Angeles(加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文从Fisher-Rao信息几何视角分析合成数据训练动态,推导出稳定且不随维度退化的收缩与不变性界,从而重新确定了防止模型崩溃所需的人类数据与合成数据最小比例。
AI中文摘要:
大型语言模型(LLMs)现在通常使用合成数据进行训练,因为高质量的人类数据已被日益增长的大规模模型需求所耗尽。然而,递归地使用合成数据进行训练常常引发模型崩溃,这是一种退化的反馈循环,模型会逐渐遗忘真实的数据分布。将合成数据与新鲜人类数据混合训练是一种合理的对策,可以防止模型崩溃。然而,一个悬而未决的问题是,为维持训练稳定性,人类数据与合成数据所需的最小比例究竟是多少。在本文中,我们为防止模型崩溃所需的最小人类数据比例建立了严格的理论保证。尽管先前的工作为该比例确立了形式化的下界,但在高维情况下该下界可能失效,因为其分析依赖于R^n中通常的欧几里得度量,并不适用于分类概率分布空间。相反,在本文中,我们通过分析Fisher-Rao度量下过程的动态,明确利用了概率单纯形的信息几何结构。我们推导了定量的收缩与不变性界,这些界是稳定的,且不会随着维度增加而变得平凡。因此,我们表明,防止模型崩溃所需的有效数据比例与先前所暗示的不同。
英文摘要:
Large Language Models (LLMs) are now routinely trained using synthetic data, since high-quality human data has been exhausted by the ever increasing needs of larger and larger models. However, recursive training on synthetic data frequently induces model collapse, a degenerative feedback loop where models progressively forget the true underlying data distribution. Training on a mixture of synthetic and fresh human data is a logical countermeasure and can prevent model collapse. However, it is an open question as to what is the exact minimum required ratio of human-to-synthetic data to maintain training stability. In this paper, we establish rigorous theoretical guarantees on the minimum rate of human data required to prevent model collapse. Although previous work established a formal lower bound for this ratio, such bound can be vacuous for very high dimensions, as the analysis relies on the usual Euclidean metric in R^n and is not adapted to the space of categorical probability distributions. Instead, in this paper we explicitly leverage the information-geometric structure of the probability simplex by analyzing the dynamics of the process under the Fisher-Rao metric. We derive quantitative contraction and invariance bounds that are stable and do not become trivial as the dimensions increase. Thus, we show that the effective required data ratio to prevent model collapse is different than previously implied.