公平性崩溃现象:在合成数据上训练的语言模型中的偏差放大
The Fairness Collapse Phenomenon: Bias Amplification in Language Models Trained on Synthetic Data
浏览论文内容
中文总结 AI 辅助
该研究发现,在合成数据上反复训练语言模型会出现公平性崩溃,即偏差在标准指标未明显下降时悄然放大,揭示了合成数据污染的关键风险。
中文摘要 AI 辅助
已有研究表明,在人工生成数据上训练的生成模型会出现模型崩溃,导致性能显著下降。随着合成内容越来越多地污染语言模型的训练语料库,这引发了关于在持续预训练中使用开放数据的关键担忧。尽管先前的工作已证明语言模型存在模型崩溃,但接触合成数据是会放大还是减弱预训练模型中已存在的社会偏差仍不清楚。由于已知语言模型会复制并放大人口统计刻板印象,对自身生成数据进行递归训练可能会形成一个自我强化的反馈循环,其中偏见关联在各代中逐渐增强,我们将这一假设现象称为公平性崩溃。在本研究中,我们构建了受控训练机制,使用Bias in Bios数据集让模型在合成数据上进行反复训练。实验中,我们观察到一致且令人担忧的模式:在标准语言建模指标出现显著下降之前,就已出现公平性退化。这一结果凸显了语言模型训练中合成数据污染相关的关键风险:在模型崩溃的强指标显现之前,偏差可能会悄然增大。
英文摘要
Generative models trained on artificially generated data have been shown to exhibit model collapse, resulting in significant performance degradation. As synthetic content increasingly contaminates the training corpora of language models, this raises critical concerns about the use of open data in continued pretraining. Although previous work has demonstrated model collapse in language models, it remains unclear whether exposure to synthetic data amplifies or attenuates the social biases already present in pretrained models. Because language models are known to reproduce and amplify demographic stereotypes, recursive training on self-generated data may create a self-reinforcing feedback loop in which biased associations become progressively stronger across generations. We call this hypothesized phenomenon fairness collapse. In this work, we construct controlled training regimes in which models are repeatedly trained on synthetic data using the Bias in Bios dataset. Across experiments, we observe a consistent and concerning pattern: fairness degradation emerges before substantial degradation is reflected by standard language-modeling metrics. This result highlights a critical risk associated with synthetic data contamination in language model training: bias can increase silently before strong indicators of model collapse become apparent.
发表机构
- Laboratoire Hubert Curien(于贝尔·屈里安实验室)
- CNRS(法国国家科学研究中心)
- Université Claude Bernard Lyon 1(里昂第一大学)
- Université Lumière Lyon 2(里昂第二大学)
- École Centrale de Lyon(里昂中央理工学院)
机构由 AI 辅助整理,请以论文原文为准。