arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

逃离低维重叠:通过高维稀疏解缠的多任务模型合并

Escaping Low-Dimensional Overlap: Multi-Task Model Merging via High-Dimensional Sparse Disentanglement

Yihang Zhang, Shengke Sun, Junjie Wen, Feng Zeng

arXiv 2608.25354首次发表:更新:

发表机构

Central South University; Nanjing University of Science and Technology; Hefei University of Technology(中南大学; 南京理工大学; 合肥工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多任务模型合并的任务干扰问题,本文提出基于SAEs的高维稀疏解缠合并框架,结合GR-ZOO实现选择性合并,在Qwen2.5系列模型的多任务上优于现有基线。

AI 中文摘要

模型合并为构建多任务通用模型提供了无需额外训练的高效方式,但其性能常因严重的任务干扰而下降。模型合并中的任务干扰主要源于「叠加」,即任务特定特征在参数空间中发生纠缠,这种纠缠使传统分解方法无法有效将有用的任务方向与干扰成分分离。本文提出一种基于稀疏表示的合并框架,使用稀疏自编码器(SAEs)将任务向量投影到高维稀疏特征空间,实现融合前的特征级解缠。为降低计算开销,我们进一步引入轻量级分组排序零阶优化器(GR-ZOO)以识别任务关键层进行选择性合并。在Qwen2.5-1.5B和Qwen2.5-7B上的实验表明,我们的方法在数学推理、代码生成、指令遵循和通用知识任务上,始终优于Task Arithmetic、TIES-Merge、DARE、Fisher-Merge等代表性基线及若干近期无训练合并方法。在Qwen2.5-1.5B的高度冲突四任务设置中,我们的方法较最强基线实现了2.78%的提升。

英文摘要

Model merging provides an efficient way to construct multi-task generalist models without additional training, but its performance often degrades under severe task interference. Task interference in model merging primarily stems from \textit{superposition}, where task-specific features become entangled within the parameter space. This entanglement renders conventional decomposition methods insufficient for effectively isolating useful task directions from interfering components. In this paper, we propose a sparse-representation-based merging framework that uses Sparse Autoencoders (SAEs) to project task vectors into a high-dimensional sparse feature space, enabling feature-level disentanglement before fusion. To reduce computational overhead, we further introduce a lightweight Group-Ranked Zeroth-Order Optimizer (GR-ZOO) to identify task-critical layers for selective merging. Experiments on both Qwen2.5-1.5B and Qwen2.5-7B demonstrate that our method consistently outperforms representative baselines, including Task Arithmetic, TIES-Merge, DARE, Fisher-Merge,and several recent training-free merging methods, across mathematical reasoning, code generation, instruction following, and general knowledge tasks. In a highly conflicting four-task setting on Qwen2.5-1.5B, our method further achieves a 2.78\% improvement over the strongest baseline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑