arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非维度,而是范数:无梯度语言模型权重扰动的关键因素

Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models

Taeyeong Kim, Ahhyun Kim, TaeHyeon Kim, Unggi Lee

arXiv 2608.01624首次发表:更新:

AI 中文总结

该研究探究无梯度语言模型权重扰动的关键因素,发现全权重搜索非必需,扰动范数是核心有效因素,其安全区域可跨模型与尺度迁移,优化方向聚焦于扰动强度的确定。

AI 中文摘要

将语言模型适配到某一任务不再需要训练其所有权重,一系列参数高效方法已将可训练参数数量从数十亿降至少数标量。无梯度适配方法通过随机采样权重扰动并保留得分较高的扰动,却未遵循这一趋势,仍会扰动权重张量的每个元素。由于现有方法同时改变搜索空间、扰动尺度和聚合方式,目前尚不清楚是否需要进行全权重搜索,以及更根本地,扰动的哪个特性使其有效。本文通过在固定流程中一次干预一个因素来解决该问题,在保持候选评分和投票不变的情况下,分别改变搜索维度、承载扰动的子空间及其范数。在49个模型-基准组合中,仅扰动12至16个标量的冻结帧平均比全权重搜索低1.8个准确率点,且在36个组合中表现更差。维度和基的选择均无法解释该性能。当随机帧与SVD帧的Grassmann重叠处于随机水平时,若匹配单一尺度因子,其表现与SVD帧相同;而在大尺度下,SVD方向会先崩溃。最终留存的是扰动范数,其可用范围在7个模型间的差异不超过5倍,且在单个模型内部保持稳定。因此,扰动范数是唯一存在失效模式的因素,且其安全区域可跨尺度和模型家族迁移。设计问题从“在哪个子空间进行扰动”缩小为“扰动的强度如何确定”。

英文摘要

Adapting a language model to a task no longer requires training all of its weights, and a line of parameter-efficient methods has driven the trainable count from billions down to a handful of scalars. Gradient-free adaptation, which samples random weight perturbations and keeps the ones that score well, has not followed that trajectory and still perturbs every entry of the weight tensor. It is unknown whether that full-weight search is necessary, and more fundamentally which property of a perturbation makes it work at all, because existing methods vary the search space, the perturbation scale, and the aggregation together. We resolve this by intervening on one factor at a time inside a fixed pipeline, holding candidate scoring and voting constant while we vary the search dimension, the subspace that carries the perturbation, and its norm. Perturbing a frozen frame of 12 to 16 scalars stays 1.8 accuracy points behind full-weight search on average across 49 model-benchmark cells, trailing it in 36 of them. Neither the dimension nor the choice of basis explains that performance. A random frame whose Grassmann overlap with the SVD frame is at chance level performs identically once a single scale factor is matched, and at large scales the SVD directions collapse first. What survives is the perturbation norm, whose usable range closes within a factor of five across seven models and stays flat inside. The perturbation norm is therefore the one factor with a failure mode, and its safe region transfers across scale and family. The design question narrows from which subspace to perturb to how hard to shake.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑