arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

预测与修复大型语言模型中的合并崩溃

Predicting and Repairing Merge Collapse in Large Language Models

Jungseob Lee, Seungyoon Lee, Sugyeong Eo, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim

arXiv 2610.03199首次发表:更新:

发表机构

Korea University; Yonsei University Mirae Campus; Konkuk University(高丽大学; 延世大学未来校区; 建国大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于任务向量方差的统计量来预测大型语言模型合并崩溃,并引入PRISM操作符通过软阈值修复破坏性合并,实验验证了其有效性。

AI 中文摘要

从共享基础模型微调而来的大型语言模型可以通过平均其任务向量进行合并,但某些合并结果远低于基础模型,且常见的合并操作符在评估前不提供任何警告。我们表明,专家任务向量的一个统计量既能预测这种崩溃,又能校准其修复。平均操作所移除的功率等于任务向量在专家间的方差,即我们定义的干扰度量。在一个有效噪声模型下,合并注入的扰动随合并系数和干扰的增加而增大,从而产生合并前得分。在我们对四个模型家族的二十二种合并配置的实验中,只有破坏性合并超过该得分的阈值。我们发现,专家间符号冲突的统计量(现有合并操作符的常见目标)具有反预测性。随后,我们在评估前预测了十四种合并的结果,其中十二种预测正确,包括一对专家因持续预训练而超过阈值并导致破坏性结果的情况。为解决这种崩溃,我们引入了PRISM,一种先平均任务向量,然后根据每层的干扰设定水平对每层进行软阈值处理的操作符。无需数据或调优,PRISM使所有五种破坏性合并保持在阈值以上,且评估噪声在基础模型范围内,而普通平均则至少低于基础模型14.4个点或完全崩溃。我们仅在阈值以上应用PRISM,对阈值以下的合并(包括所有十五种无害合并)保留普通平均。代码可在该https URL获取。

英文摘要

Large language models fine-tuned from a shared base can be merged by averaging their task vectors, but some merges collapse far below the base model, and common merge operators give no warning before evaluation. We show that one statistic of the specialists' task vectors both predicts this collapse and calibrates its repair. The power that averaging removes equals the variance of the task vectors across specialists, our measure of interference. Under a working noise model, the disturbance that a merge injects grows with the merge coefficient and with interference, yielding a pre-merge score. In our experiments on twenty-two merge configurations from four model families, only destructive merges exceed a threshold on this score. We find that statistics of sign conflict between specialists, a common target of existing merge operators, are anti-predictive. We then predicted the outcomes of fourteen merges before evaluating them, and twelve predictions were correct, including the destructive outcome of a specialist pair pushed past the threshold by continued pretraining. To address this collapse, we introduce PRISM, an operator that averages the task vectors first and then soft-thresholds each layer at a level set by the layer's interference. Without data or tuning, PRISM keeps all five destructive merges above the threshold within evaluation noise of the base model, where plain averaging falls at least 14.4 points below it or collapses entirely. We apply PRISM only above the threshold and keep the plain average for merges below it, which include all fifteen harmless ones. Code is available at https://github.com/js-lee-AI/PRISM.

Comments23 pages, 5 figures, 20 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑