当模型合并与联合多任务强化学习相抗衡时:任务向量几何分析
When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis
- Aimpoint Digital Labs(Aimpoint数字实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究在强化学习中模型合并能否替代联合多任务训练,通过训练Qwen3 - 8B专家并合并,与联合训练模型对比,发现任务向量几何结构中方向和支持解耦,合并效果与联合训练相当,还发布了代码和统计数据。
AI中文摘要:
模型合并被推广为联合多任务训练的替代方案,但在强化学习环境中,这种替代从未与它声称要取代的基线进行过测试。我们通过在AppWorld智能体基准上使用LOOP训练难度1和难度2的Qwen3 - 8B专家,然后合并它们并与在相同数据上联合训练的模型进行比较。在任务目标完成方面,合并与联合强化学习效果相当。为解释为何合并方法在此无关紧要,我们测量了专家任务向量的几何结构,发现其方向和支持是解耦的,基于支持和符号的合并会归结为近似均匀平均。我们还发布了所有代码和统计数据。
英文摘要:
Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting this substitution is essentially never tested against the baseline it claims to replace: methods merge independently released agents precisely because a joint model is unavailable. We build the missing comparison. Training difficulty-1 and difficulty-2 Qwen3-8B specialists on the AppWorld agent benchmark with LOOP, we merge them (TIES, RAM+) and pit the result against a jointly trained model on the same data. On task-goal completion, merging matches joint RL -- and every merge variant is statistically indistinguishable. To explain why merge method does not matter here, we measure the geometry of the specialists' task vectors, which carries no task-sampling noise: they are near-orthogonal (cosine 0.06 - 0.10) despite ~65% support overlap, a small, shared direction that grows over training and that we calibrate against a random-init floor and a same-run ceiling to confirm it reflects learning, not the low-rank parameterization. Because direction and support are decoupled, support and sign-based merging (RAM, TIES) collapse to near-uniform averaging. We release all code and statistics.