arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40295cs.CLcs.LG

一个AI代币值多少钱?野生AI生成网页文本的缩放定律

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

  • University of Maryland(马里兰大学)
  • Pangram Labs(Pangram 实验室)

机构由 AI 辅助整理,请以论文原文为准。

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI总结:

本研究通过预训练800个语言模型,提出一种新的缩放定律,揭示了野生AI生成文本对预训练的影响:初期有益但随后有害,并提供了过滤和重复人类文本的建议,同时发布了WildAI语料库。

AI中文摘要:

网页文本构成了预训练数据的主体,并且日益由AI生成。在应用FineWeb质量过滤后,我们发现2026年6月网页数据中27.5%的token被Pangram标记为AI生成,到8月这一比例上升至31.1%。与合成数据或模型崩溃设置不同,这种“野生”AI文本来自众多模型,面向人类读者编写,并以未标记的形式进入预训练语料库。野生AI文本如何影响语言模型的预训练?为了回答这个问题,我们预训练了800个语言模型,改变AI token与人类token的添加比例,并拟合缩放定律以预测在人类和AI生成文本上的保留损失。对于数据匮乏的模型,在预训练数据中添加AI token最初会降低人类文本的损失,但随着添加更多,收益饱和并迅速“逆转”为损害。对于在高预算人类文本上训练的模型,AI token几乎立即提高损失,而同样数量的新鲜人类token则持续降低损失。Hoffman等人(2022)的缩放定律未能预测这种行为。我们提出了一种新的缩放定律,包含独立的收益和损害项,允许AI token的价值改变符号,同时在缺乏AI文本时简化为Chinchilla。当在较小模型上拟合时,我们的缩放定律预测了AI文本对保留人类文本损失的影响,对于比其大3.6倍的模型,在所有AI比例下,误差比最佳现有定律低41%。我们建议在目标是人类文本时过滤AI文本,在扩展训练数据集时优先重复人类文本,并分别报告人类和AI文本的验证损失;当目标是AI文本时,AI文本仍然有价值。我们发布了WildAI,一个包含83B token的语料库,带有AI、主题和格式标签,以及所有800个模型和代码,网址为https URL。

英文摘要:

Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models, varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly *reverses* into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it. Scaling laws such as Hoffman et al. (2022) fail to predict this behavior. We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios. We recommend filtering AI text when the target is human text, repeating human text before expanding the training dataset with AI-generated web text, and reporting validation loss on human and AI text separately AI text remains valuable when the target is AI text. We release WildAI, an 83B-token corpus with AI, topic, and format labels, all 800 models and code at https://github.com/pangramlabs/WildAI.

↑