何时微调的一阶模型能够界定遗忘?
When Can First-Order Models of Fine-Tuning Bound Forgetting?
浏览论文内容
中文总结 AI 辅助
本研究探究微调一阶模型能否界定遗忘,发现仅当事实的响应系数变化小于其到边界距离时,完整Freedman界可可靠界定遗忘概率,并经预注册实验验证。
中文摘要 AI 辅助
对语言模型进行新数据微调可能会使其遗忘本应保留的事实。我们探究,在微调运行开始时进行的测量,能否针对每个受保护事实,界定该运行使模型遗忘该事实的概率。在对参数规模从0.6B到14B的模型进行LoRA微调并采用随机梯度下降时,通过有限差分探针估计的一阶响应模型,预测每个事实边际变化的相关系数在0.974至0.998之间。然而,基于该模型构建的遗忘预测却失败了,因为遗忘所需的参数变化远远超出了模型得到验证的区域。不过,探针能够界定边际首次降至接近零的边界以下的概率:我们为边际的线性替代量推导了Freedman和Azuma首达界,并在新的运行中测试这些界是否对模型成立。这些界包含一个项R,用于衡量运行期间响应系数的变化程度。简化版Freedman界(设R=0)认证了大多数事实,但在112种条件中的14种情况下被违反,且每个被违反的事实都有R≥a,其中a是该事实边际到边界的距离。完整版Freedman界仅认证满足R<a的事实,且在所有条件下均成立。在被违反的事实上,边际在测试运行间的离散度中位数是响应模型预测值的14.6倍,因此这些失败是模型的失效,而在我们的数据中,它们仅发生在R≥a的情况下。我们发现这一模式是事后性的,并在两项预注册的验证性研究中对其进行了测试,共包含43个新条件:完整版界在所有条件下均成立,而简化版界仅在3个事实上失败,且每个都有R≥a。因此,对于响应系数变化小于其到边界距离的事实,微调的一阶模型能够界定遗忘概率。
英文摘要
Fine-tuning a language model on new data can make it forget facts that it should keep. We ask whether measurements taken at the start of a fine-tuning run can bound, for each protected fact, the probability that the run makes the model forget it. In LoRA fine-tuning with stochastic gradient descent on models from 0.6B to 14B parameters, a first-order response model estimated by finite-difference probes predicts changes of per-fact margins with correlation 0.974-0.998. Predictions of forgetting built on this model nevertheless failed, because forgetting requires parameter changes far outside the region in which the model was validated. The probes can, however, bound the probability that a margin first falls below a boundary near zero: we derive Freedman and Azuma first-passage bounds for a linear surrogate of the margin and test on new runs whether they hold for the model. The bounds contain a term R that measures how much the response coefficients change during the run. The simplified Freedman bound, which sets R = 0, certified most facts but was violated in 14 of 112 conditions, and every fact on which it was violated had R >= a, where a is the distance of the fact's margin to the boundary. The complete Freedman bound certifies only facts with R < a, and it held in every condition. On the violated facts, the spread of the margin across test runs was a median of 14.6 times the prediction of the response model, so the failures are breakdowns of the model, and in our data they occurred only where R >= a. We found this pattern post hoc and tested it in two preregistered confirmatory studies with 43 new conditions: the complete bound held in all of them, and the simplified bound failed there on only 3 facts, each with R >= a. First-order models of fine-tuning can thus bound forgetting on the facts whose response coefficients change by less than their distance to the boundary.