AI 中文总结
提出MeRoTune,利用与RoPE可交换的缩放旋转矩阵类,通过可调混合比例学习校正矩阵,实现安全且可事后调整的模型合并。
AI 中文摘要
当你通过简单地对权重求平均来合并来自同一基础检查点的两个微调模型时,你隐含地假设它们的注意力子空间仍然对齐。最近的工作尝试通过学习每个模型的查询和键投影的可逆校正矩阵$M$来修复错位。这种校正会在点积之前抵消——在查询侧使用$M$,在键侧使用$M^{-T}$。然而,只有当投影和点积之间没有任何东西时,这种抵消才是精确的。实际上,几乎所有现代开放权重语言模型都在那里放置了旋转位置嵌入(RoPE)。在本文中,我们证明了在RoPE下,这种抵消是精确的当且仅当$M$与RoPE的每位置旋转可交换。我们推导出满足此条件的特定矩阵类:在每个RoPE频率对内独立作用的缩放旋转。这构成了当前方法通常训练的无约束矩阵的一个严格的低维子集。基于此,我们将这个受约束的矩阵类转化为一种新的合并方法。在保持基础权重完全冻结的同时,两个微调模型各自学习自己的符合RoPE的校正矩阵。我们针对选定的混合比例优化这些校正,以便最终结果可以像旋钮一样事后调整,而不是锁定在单一的固定合并中。我们的默认方法在固定的混合比例下训练,类似于LoRA预先设置其缩放超参数的方式。我们还尝试在每个训练步骤随机重新采样混合比例,并报告两种方法的结果。
英文摘要
When you merge two fine-tuned models from the same base checkpoint by simply averaging their weights, you implicitly assume their attention subspaces are still aligned. Recent work attempts to fix misalignments by learning an invertible correction matrix, $M$, for each model's query and key projections. This correction cancels out---using $M$ on the query side and $M^{-T}$ on the key side---right before the dot product. However, this cancellation is only exact if nothing sits between the projection and the dot product. In reality, almost all modern open-weight language models put a rotary position embedding (RoPE) exactly there. In this paper, we show that this cancellation is exact under RoPE if and only if $M$ commutes with RoPE's per-position rotation. We derive the specific class of matrices where this holds: a scaled rotation acting independently within each RoPE frequency pair. This forms a strict, low-dimensional subset of the unconstrained matrices that current methods normally train. Building on this, we turn this constrained matrix class into a new merging method. While keeping the base weights entirely frozen, two fine-tunes each learn their own RoPE-compliant correction matrices. We optimize these corrections against a chosen blend ratio so the final result can be adjusted post-hoc like a dial, rather than locked into a single fixed merge. Our default approach trains at one fixed blend ratio, similar to how LoRA sets its scaling hyperparameter in advance. We also experiment with resampling the blend ratio randomly at every training step, and we report the results of both approaches.