发表机构
McGill University; Mila – Quebec AI Institute; Université de Montréal; Nanyang Technological University(麦吉尔大学; 魁北克人工智能研究所; 蒙特利尔大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种基于热带几何的门控感知逐单元投影器,通过精确保护开放词元子空间,在微调中显著减少遗忘,并在多个模型上优于现有方法。
AI 中文摘要
在语言模型上对新文本进行微调会降低其已有的能力。诸如Adam-NSCL和GPM等无重放投影器在更新的每一行中禁止使用层输入的某个共享子空间。ReLU层的热带几何解释了为何这种方法过于粗糙。在数据空间中,单元的边界是热带超曲面,其胞腔与一个zonotope的上顶点对偶;在权重空间中,每个旧词元是一个超平面,这些词元切割出一个多面体,即保持每个词元在其正确一侧的权重的闭包。一个精确的恒等式将这两个图景联系起来:在任意权重变化下层输出的平方变化分解为胞内、开放到封闭和封闭到开放三项,其中前两项作用于每个单元激活的词元上。该恒等式命名了一个门控感知的逐单元投影器,并且一个预算分离定理为精确保护定价:它花费一个单元其自身开放词元的秩,而一个共享子空间在每一行至少支付其并集的秩。在OPT-1.3b上,其中96%的(词元,单元)对是封闭的,该投影器在从每行9到60个约束方向的所有六个匹配预算下比Adam-NSCL遗忘得更少,差距从1.1倍扩大到4.3倍;在GPM的能量阈值下,使用其1/5.5的方向数,它将Adam-NSCL的遗忘减半。在OPT-6.7b上,它在匹配预算下匹配了Adam-NSCL的遗忘,同时学习了更多。正如理论所预测的,开放/封闭划分是操作变量:开放词元在18个种子对中的18个上优于随机、符号盲和反门控词元集。在剪枝修复中,每个基于导数的稠密权重输出误差局部模型对开放的词元对视而不见:门控加权目标的最小化器可以离开多面体,该目标的闭式解在OPT-1.3b上比不修复差1.94 nats,而一个凸的单侧惩罚限制了逃逸。
英文摘要
Fine-tuning a language model on new text degrades what it already does. Replay-free projectors such as Adam-NSCL and GPM forbid one shared subspace of a layer's inputs in every row of the update. The tropical geometry of a ReLU layer shows why this is too coarse. In data space, the units' walls are tropical hypersurfaces whose cells are dual to the upper vertices of a zonotope; in weight space, each old token is a hyperplane, and the tokens cut out a polyhedron, the closure of the weights that keep every token on its side. An exact identity joins the two pictures: the squared change of the layer's output under any weight change splits into in-cell, open-to-closed and closed-to-open terms, and the first two live on the tokens each unit fires on. The identity names a gate-aware per-unit projector, and a budget-separation theorem prices exact protection: it costs a unit the rank of its own open tokens, while a shared subspace pays at least the rank of their union in every row. On OPT-1.3b, where 96% of (token, unit) pairs are closed, the projector forgets less than Adam-NSCL at all six matched budgets from 9 to 60 constrained directions per row, the gap widening from $1.1\times$ to $4.3\times$; with 1/5.5 of the directions it halves the forgetting of Adam-NSCL at GPM's energy threshold. On OPT-6.7b, it matches Adam-NSCL's forgetting at matched budget while learning more. As the theory predicts, the open/closed partition is the operative variable: open tokens beat random, sign-blind and anti-gate token sets on 18 of 18 seed-pairs. In pruning repair, every derivative-based local model of the output error at the dense weights is blind to pairs that open: the minimisers of the gate-weighted objective can leave the polyhedron, the objective's closed-form solution is 1.94 nats worse than no repair on OPT-1.3b, and a convex one-sided penalty bounds the escape.