预训练数据投毒的可解理论:状态依赖的标度指数
A Solvable Theory of Pre-training Data Poisoning: Regime-Dependent Scaling Exponents
AI总结:
该论文提出可解理论,证明预训练数据投毒导致的性能退化指数取决于极限顺序,并解释非整数标度指数源于重尾奇异结构。
AI中文摘要:
大语言模型的预训练数据投毒通常通过针对性后门及其在安全后训练中的存留来研究,这留下了一个更基本的问题:随着投毒率 $\varepsilon$ 增长,模型的干净数据性能如何退化?受我们对OLMo风格模型的受控预训练运行的启发,在这些运行中,架构、令牌预算和优化计划匹配的投毒模型与干净模型之间的相对干净数据验证困惑度增加 $\Delta$ 很好地拟合为幂律 $\Delta\approx C\varepsilon^{a}$,且指数为非整数,我们询问这样的定律在理论上需要什么条件。我们首先证明一个解析性障碍:每当受污染目标在非退化干净数据最优点附近对 $\varepsilon$ 解析依赖时,$\Delta$ 通常对 $\varepsilon$ 是二次的,因此一般的非整数指数是真正奇异结构的标志。然后,我们在可解的截断岭回归中提供该结构,该回归具有重尾协变量,由 $q_\star$ 控制,以及标签偏移投毒。我们的核心结果是,超额风险标度指数取决于极限的顺序:在高维比例状态下为 $\epsilon^{q_\star/(q_\star+2)}$,而先取充足数据极限则给出 $\epsilon^{2-2/q_\star}$,且极限不可交换。我们通过多个数值模拟确认了这一预测。最后,我们认为有限训练时间在局部LLM预训练中扮演曲率谱截断的角色,在重尾逆曲率谱下推导出观察到的标度律作为建模假设。
英文摘要:
Pre-training data poisoning of large language models is usually studied using targeted backdoors and their survival through safety post-training, which leaves open a more basic question: how does a model's clean data performance degrade as the poison rate $\varepsilon$ grows? Motivated by our controlled pre-training runs of OLMo-style models, in which the relative clean data validation perplexity increase $Δ$ between poisoned and clean models matched in architecture, token budget, and optimization schedule is well fit by a power law $Δ\approx C\varepsilon^{a}$ with a non-integer exponent, we ask what such a law requires theoretically. We first prove an analyticity barrier: whenever the contaminated objective depends analytically on $\varepsilon$ around a nondegenerate clean data optimum, $Δ$ is generically quadratic in $\varepsilon$, so a generic non-integer exponent is a signature of genuinely singular structure. We then supply that structure in solvable truncated ridge regression with heavy-tailed covariates, controlled by $q_\star$, and a label-shift poisoning. Our central result is that the excess risk scaling exponent depending on the order of limits: in the higher dimensional proportional regime it is $ε^{q_\star/(q_\star+2)}$, whereas taking the ample-data limit first gives $ε^{2-2/q_\star}$, and the limits do not commute. We confirm this prediction through several numerical simulations. Finally, we argue that finite training time plays the role of truncation on the curvature spectrum in local LLM pre-training, deriving the observed scaling law under heavy tailed inverse curvature spectrum as a modeling hypothesis.