任务向量何时产生干扰?映射权重空间组合的有效性边界
When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space Composition
AI总结:
该研究探究任务向量在权重空间组合的有效性边界,通过Qwen2.5、Llama-3.1等模型实验,发现其支持粗略的功能陈述而非通用合并性能预测器。
AI中文摘要:
任务算术将微调位移视为权重空间中可组合的方向,但参数加法何时能反映模型函数的可预测变化仍不清楚。我们将参数几何与功能几何分离,在二维任务向量表面上测量成对功能非加性,使用基于输入分布的首-token预测分布交互比率,结合范数匹配对照、三个训练种子及仅响应微调进行评估。在Qwen2.5-1.5B上,代码+安全任务在代码和指令提示下比匹配的代码+数学对照更具非加性,但在数学提示下并非如此。在预先指定的六任务扩展中,所有8对未见任务对的高低比较均符合预测符号,该主要排序在0.5B全参数微调、Qwen2.5 LoRA规模测试(最高7B)及Llama-3.1-8B跨架构审计下仍持续存在。外部验证揭示了更清晰的边界:原始公开代码、指令及安全提示保留了连续对比,而指令风格包装在相同公开代码提示上使对比消失,EvalPlus pass@1交互则未稳健复现该现象。因此,权重空间组合支持跨适配方法、规模及一个额外模型家族的粗略、输入和格式相关功能陈述,而非通用的合并性能预测器。
英文摘要:
Task arithmetic composes skills by adding weight displacements, and merged models are then judged on benchmark suites. We measure when that composition is functionally additive, and find that the answer depends as much on how the model is prompted as on which tasks are merged. Across two-dimensional composition surfaces -- five model settings from 0.5B to 8B, two families, LoRA and full fine-tuning -- pairwise non-additivity is real, seed-stable, and transfers in coarse order to unseen task pairs: all eight preregistered sign predictions held. But it is input-conditioned everywhere we measured: the same merged model that shows a six-point interaction contrast on code prompts shows none on math prompts, and wrapping the identical code prompts in the instruction template the adapters were trained on collapses the contrast twenty-fold, from +6.9 to +0.3 points -- while re-serializing them in an untrained chat template leaves it intact (+12.5), falsifying our own preregistered prediction. Execution benchmarks (pass@1) inherit the training-format wrapper's blindness. Weight-space composition therefore supports coarse, input- and format-conditioned functional statements -- not a universal merging-performance predictor, and not one that training-format evaluations can see.