Omega-S:一种用于大语言模型微调的功能弹性指数
Omega-S: A Functional Resilience Index for LLM Fine-Tuning
AI总结:
Omega-S是一种仅依赖权重矩阵的即插即用惩罚项,在Llama-3-8B的LoRA微调中,它在保留模型原有能力上优于无正则化、调优后的权重衰减和EWC,且成本增加不到4%。
AI中文摘要:
在新数据上微调大语言模型会导致其先前学习到的知识发生退化。我们提出了Omega-S,这是一种仅从权重矩阵计算的即插即用惩罚项:它不需要先前任务的数据、Fisher矩阵,也不需要存储旧权重的副本。它只需在现有训练循环中添加三行代码,且单步训练的成本增加不到4%。保留能力:在Llama-3-8B模型上使用LoRA,从代码微调至散文,通过HumanEval在10个随机种子上进行测量,Omega-S在10个种子中有9个保留了比无正则化更多的原始能力(绝对pass@1从0.173提升至0.238;单侧符号检验p=0.011,Wilcoxon检验p=0.006),保留率从62.9%提升至84.1%。它还在10个种子上优于调优后的权重衰减(p=0.002),在8个种子上优于调优后的EWC(p=0.014),所有对比项均在同一会话中重新测量。机制:Omega-S是拓扑结构的,其目标函数基于Tr(A^3),但我们测量了它的四个因子中实际发生变化的有三个,它们相对于权重的弹性在1e-4或更低,而度方差项的弹性为9e-3。实现后,该复合项简化为对节点度方差的惩罚,即方形模块中的行幅度和非方形模块中的方向对齐。我们报告这一点,是因为一个名称承诺一种功能但梯度表现为另一种功能的方法应当明确说明。我们还列举了开放的设计选择,包括一种对比度保持的构造,该构造实现了其设计目标,但在所有10个种子上使保留率变差。重复相同配置、相同种子和相同硬件,保留率的标准差为0.104。我们未发现文献中对语言模型的低秩微调有此量化结果,它限定了该领域中每一对种子配对的比较,包括我们的研究。代码、每个种子的结果以及所有负面结果的完整记录均可用。
英文摘要:
Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S, a drop-in penalty computed from the weight matrix alone: it needs no previous-task data, no Fisher matrix and no stored copy of the old weights. It is three lines in an existing training loop and adds under 4% to the cost of a step. Retention. On Llama-3-8B with LoRA, fine-tuned from code to prose and measured by HumanEval over ten seeds, Omega-S retains more of the original capability than no regularisation on 9 of 10 seeds (0.173 -> 0.238 absolute pass@1; sign test one-sided p=0.011, Wilcoxon p=0.006), as a retention ratio, 62.9% -> 84.1%. It also beats tuned weight decay on 10 of 10 seeds (p=0.002) and tuned EWC on 8 of 10 (p=0.014), every arm re-measured in the same session. Mechanism, measured rather than asserted. Omega-S is topological by construction, its objective built from Tr(A^3), but we measured which of its four factors actually moves and three do not: their elasticity with respect to the weights is at or below 1e-4, against 9e-3 for the degree-variance term. As implemented, the composite reduces to a penalty on the variance of node degrees, which means row magnitude in square modules and directional alignment in non-square ones. We report this because a method whose name promises one thing and whose gradient does another should say so. We also enumerate the open design choices, including a contrast-preserving construction that does what it was designed to do and makes retention worse on all ten seeds. Repeating an identical configuration, same seed and same hardware, gives a standard deviation of 0.104 in retention ratio. We have not found this quantified for low-rank fine-tuning of language models, and it bounds every seed-paired comparison in this literature, ours included. Code, per-seed results and the full record of negative results are available.