一阶可预测但成对脆弱:训练后变压器中的局部任务适应
First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers
- DAIMLD(数据挖掘与机器学习系)
- St. Petersburg Department of the Steklov Institute of Mathematics(斯捷克洛夫数学研究所圣彼得堡分部)
- St. Petersburg State University(圣彼得堡国立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究训练后变压器中局部任务适应,通过多任务LoRA操作点测量相关属性,发现一阶变化可预测,成对结构脆弱,如两更新顺序敏感、任务梯度子空间旋转等,还给出连续任务梯度步骤相关项及预测与起始尺度情况。
AI中文摘要:
任务算术、顺序微调、激活引导和一阶随机搜索都通过围绕已训练检查点的相对小扰动进行操作,且依赖不同局部近似。我们围绕多任务LoRA操作点,在9个变压器(82M - 7B)上,用前瞻性注册的属性列表、阈值和测试分割来测量8个此类属性。发现共享的单向有效性窗口直至测试尺度\(10^{-2}\),但成对组合或更新排序无通用半径。单个方向上探针损失变化在整个网格中保持一阶可预测,而成对结构更脆弱,在超三分之一测量组合中,两更新顺序敏感性在窗口内严格出现,任务梯度子空间在几十步内旋转等。对于两个连续任务梯度步骤,主导阶相关项是李括号\(H_B\textbf{g}_A - H_A\textbf{g}_B\),其归一化预测\(c(\eta)=\eta\kappa + O(\eta^2)\)以中位数比率1.002跟踪测量缺陷,起始尺度\(\eta^\dagger\approx0.10/\kappa\)跨模型和任务对跨越三个数量级。
英文摘要:
Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space. We measure 8 such properties with the same harness around a multitask LoRA operating point, on 9 transformers (82M-7B), with a prospectively registered property list, thresholds, and test split. We find a shared one-direction validity window up to the tested scale $10^{-2}$, but no universal radius for pairwise composition or update ordering. Along individual directions, changes of the probe loss remain first-order predictable throughout the grid: a perturbation's effect on the loss is essentially its projection onto the gradient, which is also what makes local random search work. Pairwise structure, however, proves to be far more fragile: on over a third of the measured (model, task pair) combinations, two-update order sensitivity sets in strictly inside that window; task-gradient subspaces rotate within tens of steps; additivity under our fixed activation probe fails at full task-vector scale on several models, including both held-out 7B models; and no model median passes the registered global mean-vector weight-to-steering correspondence bar. For two sequential task-gradient steps, the leading order-dependent term is the Lie bracket $H_B\textbf{g}_A-H_A\textbf{g}_B$; its normalized prediction $c(η)=ηκ+O(η^2)$ tracks the measured defect at median ratio 1.002, while the onset scale $η^\dagger\approx0.10/κ$ spans three orders of magnitude across models and task pairs.