发表机构
UNSW Sydney(新南威尔士大学悉尼分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现合并语言模型的任务向量干扰核心是方向而非幅度,沿该因果方向擦除可剂量依赖性消除干扰,且预先注册了所有46项预测并如实报告证伪结果。
AI 中文摘要
基于任务算术的模型合并技术曾一度有效,而该领域此前将问题归因于幅度因素:分层表示偏差、跨任务线性度偏差、参数重叠。本文通过因子分类账追踪合并大语言模型(LLM)的精确分层交叉项并直接对其进行干预,发现幅度作为诊断轴是不足的,且在不同模型家族间不一致。对分层通量的精确分解显示,其主要由现有交叉项的放大传输主导(在两个模型家族中占比约65%-70%,在后期模块增益大于1);擦除该交叉项会被传播抵消——其范数会重建至99%,余弦相似度达0.99——除非在输出附近应用擦除;采用6种初始位移的 basin 测试证实,被传输的方向是前向传播的吸引子。该方向具有因果载荷:沿该方向擦除会剂量依赖性地消除已表达的干扰,且在完全擦除时达到饱和,而范数匹配的反向控制则失效或产生反效果。指令包装器会调控该效应:在内部放大交叉项的包装器下,相同擦除操作需移除的相对干扰减少13倍,因为该包装器将交互淹没在模板固定的主效应中而非缩小它——该结构在其他指令模板中可复现,但在长度匹配的对照中无法复现。相比之下,幅度最多只是一个粗略的关联,而 naive bfloat16 生成的惊人±15%「普遍性」实则是量化粗糙度导致的。局部交叉项生成差异最多为1.9倍的任务对,其可因果移除的干扰差异达14倍至337倍。所有46项预测均已预先注册并在获取数据前冻结;包括我们自身头条预期及经验证的连续端点下行为恢复在内的证伪结果均被如实报告。
英文摘要
Task-arithmetic merging works until it doesn't, and the field diagnoses why by measuring interference inside the merged model. We take the most direct such measure, the exact layerwise activation cross-term of a factorial ledger, establish its causal anatomy, and then ask what it tracks. The anatomy is clean: each block mostly transports and amplifies the cross-term rather than generating it; erased, it is regenerated by the untouched marginal paths to 99% of its norm unless removed late; its output effect varies monotonically with the displacement's angle (orthogonal displacements make interference worse), and a two-assumption model derives the angle law and retro-dicts the dose curve (R^2 >= 0.99). What the measure tracks is not what the field assumes. Behavioural expert-likeness is decoupled from it across four instruments. Its cross-condition behaviour is denominator-dominated: an instruction template pins the main effect to within 1% while the absolute interaction grows 111x from two to six merged tasks, suppressing expressed interference at k=2 and amplifying it at k=6. And where merging actually collapses, the cross-term is a bystander, not the carrier: across two collapse parameterizations at two scales, even erased persistently at every position, removing it entirely repairs none of the collapse. There the output-side ratio carries no method information under a common counterfactual, while two state-space measures the field already uses rank methods correctly at both scales. All 81 predictions were frozen before their data; falsifications are reported as such. Output-side interference measurement reads the gate, the denominator, and the displacement budget, not the interference. What fails a merge is the carrier-bystander split: collapse rides in the marginal displacements while the cross-term merely accompanies it, and only state space sees the carrier.