发表机构
Iowa State University; BRAC University(爱荷华州立大学; BRAC大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过受控手术干预,证明解码器专用语言模型中集中上谱尾结构对推理任务具有计算性作用,并提出乘积目标分解方法,在多个基准上验证其破坏性及与行为转变的关联。
AI 中文摘要
权重空间结构常与语言模型行为相关,但仅凭相关性并不能确立计算上的参与性。我们通过受控干预研究解码器专用Transformer中集中出现的上谱尾。在固定的相对偏移下,我们推导出一个有限宽度条件界,将平方奇异值的逆参与比与中心softmax前logit峰度联系起来。随后,我们定义了一个逐点查询-键(QK)乘积尾目标,并将独立因子手术与保留原生注意力计算的乘积目标分解进行比较。在三个基础检查点和五个推理基准上,外加一个单独分析的指令微调Phi检查点,学习到的尾编辑在所有20个模型-任务单元中比五个固定谱匹配的Haar对照的平均值更具破坏性。十八个配对对比在Holm校正后仍然显著,而两个仅具有方向性但无定论。乘积目标因子获得更高的保留尾子空间分数,为乘积级和因子级干预之间提供了经验桥梁。组件隔离识别出QK、值-输出和多层感知机块的贡献,尽管定理仅涵盖QK。在单独的研究中,逆参与比在匹配交叉规则下先于合并准确率转变,且残差化的尾感知低秩自适应(LoRA)比标准LoRA和PiSSA更早达到目标,而最终分数区间重叠。结论仅限于所评估的检查点、层、任务、干预和对照。
英文摘要
Weight-space structure often correlates with language-model behavior, but correlation alone does not establish computational involvement. We study concentrated upper spectral tails in decoder-only transformers through controlled interventions. At a fixed relative offset, we derive a finite-width conditional bound linking the inverse participation ratio of squared singular values to central pre-softmax logit kurtosis. We then define a pointwise query--key ($QK$) product-tail target and compare independent factor surgery with a product-targeted factorization that preserves native attention computation. Across three base checkpoints and five reasoning benchmarks, plus an instruction-tuned Phi checkpoint analyzed separately, the learned-tail edit is more damaging than the mean of five fixed spectrum-matched Haar controls in all 20 model--task cells. Eighteen paired contrasts remain significant after Holm correction, while two are directional but inconclusive. Product-targeted factors attain higher held-out tail-subspace fractions, providing an empirical bridge between product- and factor-level interventions. Component isolation identifies contributions from $QK$, value--output, and multilayer-perceptron blocks, although the theorem covers only $QK$. In separate studies, inverse participation precedes pooled accuracy transitions under a matched crossing rule, and residualized tail-aware low-rank adaptation (LoRA) reaches targets earlier than standard LoRA and PiSSA while final-score intervals overlap. Conclusions are restricted to the evaluated checkpoints, layers, tasks, interventions, and controls.