发表机构
Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对异构视觉模型合并难题,提出黎曼-洛伦兹参数融合(RLPF)方法,通过双曲空间对齐与融合ViT和SSM,在多个基准上超越最佳父模型,验证了几何感知融合的潜力。
AI 中文摘要
扩展深度学习面临关键瓶颈:数据枯竭、指数级训练成本以及资源集中。模型合并无需梯度下降即可组合预训练检查点,相比重新训练可节省数个数量级的成本。当独立训练的视觉模型架构和参数形状不同时,组合它们很困难。现有的权重空间合并方法通常假设对齐的、形状兼容的检查点,而视觉Transformer(ViT)和状态空间模型(SSM)使用不同的算子实现令牌混合。我们研究了一种混合异构合并设置,该设置保留两种架构,同时按语义角色对齐参数组。我们提出的黎曼-洛伦兹参数融合(RLPF)方法将对齐的组投影到公共坐标,将选定的坐标提升到双曲空间的洛伦兹双曲面模型,计算正则化的测地线重心,并将结果解码到两个分支。然后,一个学习的门控为每个输入组合分支逻辑。组件组使用固定的曲率值,归一化参数视为欧几里得。在本手稿可用的结果中,微调系统在CIFAR-10上获得82.37%的准确率,在Oxford-IIIT Pet上获得75.04%的准确率,在ImageNet-1K上获得78.58%的top-1准确率;相应的最佳父模型准确率分别为76.54%、71.42%和76.42%。在ImageNet-1K上,报告的微调前初始化达到77.80%。这些结果支持对几何感知异构融合的进一步研究,但不支持无需训练的单检查点合并:RLPF是一个双分支混合模型,其门控和报告最终模型是经过训练的。
英文摘要
Scaling deep learning faces critical bottlenecks: data exhaustion, exponential training costs, and resource concentration. Model merging combines pre-trained checkpoints without gradient descent, offering orders-of-magnitude savings versus retraining. Combining independently trained vision models is difficult when their architectures and parameter shapes differ. Existing weight-space merging methods generally assume aligned, shape-compatible checkpoints, whereas a Vision Transformer (ViT) and a state-space model (SSM) implement token mixing with different operators. We study a hybrid Heterogeneous merging setting that retains both architectures while aligning parameter groups by semantic role. Our proposed Riemannian--Lorentz Parameter Fusion (RLPF) method projects aligned groups to common coordinates, lifts selected coordinates to the Lorentz hyperboloid model of hyperbolic space, computes a regularized geodesic barycenter, and decodes the result into the two branches. A learned gate then combines branch logits for each input. Component groups use fixed curvature values, with normalization parameters treated as Euclidean. In the results available in this manuscript, the fine-tuned system obtains 82.37\% on CIFAR-10, 75.04\% on Oxford-IIIT Pet, and 78.58\% top-1 accuracy on ImageNet-1K; the corresponding best-parent accuracies are 76.54\%, 71.42\%, and 76.42\%. On ImageNet-1K, the reported pre-fine-tuning initialization reaches 77.80\%. These results support further study of geometry-aware heterogeneous fusion, but not a training-free single-checkpoint merge: RLPF is a two-branch hybrid whose gate and reported final models are trained.