arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24612cs.CV

超越均匀子空间:面向多任务模型合并的频谱感知与深度自适应融合

Beyond Uniform Subspaces: Spectrum-Aware and Depth-Adaptive Fusion for Multi-Task Model Merging

发表机构中国科学技术大学 · 北京通用人工智能研究院 · 北京科技大学
查看机构详情
  • University of Science and Technology of China(中国科学技术大学)
  • BIGAI(北京通用人工智能研究院)
  • University of Science and Technology Beijing(北京科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Ruxi Gu, Zilei Wang, Wei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有模型合并方法忽略任务频谱与深度异质性的问题,提出SADA-Merging框架,通过频谱感知的容量分配、频谱保留和深度锚定,实现无数据多任务模型合并,在多种设置下优于现有方法。

中文摘要 AI 辅助

模型合并旨在无需额外训练过程的情况下整合多个任务特定模型。然而,现有的基于子空间的方法在很大程度上依赖于对任务更新的均匀处理,忽视了其内在的频谱和深度异质性。我们识别出该假设的两个关键偏差:不同任务需要不同的子空间容量,并且对频谱变换表现出不同的容忍度,而子空间投影会引入与深度相关的失真。基于这些观察,我们提出了SADA-Merging,一个用于无数据模型合并的频谱感知与深度自适应框架。SADA-Merging根据频谱复杂度分配任务特定的子空间容量,根据任务可塑性调整频谱保留,并应用深度相关的锚定来补偿投影引起的失真。这使得融合过程能够适应每个任务的内在几何结构及其在网络深度上的敏感性。SADA-Merging直接作用于任务更新,适用于全微调和LoRA设置。大量实验表明,在不同任务规模和适应设置下,该方法相对于现有无数据合并方法具有一致的改进。

英文摘要

Model merging aims to consolidate multiple task-specific models without access to extra training process. However, existing subspace-based methods largely rely on a uniform treatment of task updates, overlooking their intrinsic spectral and depth-wise heterogeneity. We identify two key deviations from this assumption: different tasks require different subspace capacity and exhibit different tolerance to spectral transformation, while subspace projection introduces depth-dependent distortion. Based on these observations, we propose SADA-Merging, a spectrum-aware and depth-adaptive framework for data-free model merging. SADA-Merging allocates task-specific subspace capacity according to spectral complexity, adapts spectral preservation according to task-wise plasticity, and applies depth-dependent anchoring to compensate for projection-induced distortion. This enables the fusion process to adapt to both the intrinsic geometry of each task and its sensitivity across network depth. SADA-Merging operates directly on task updates and is applicable to both full fine-tuning and LoRA settings. Extensive experiments demonstrate consistent improvements over existing data-free merging methods across different task scales and adaptation settings.

补充信息

↑