EGGROLL,展开:理解并改进大规模低秩进化策略
EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale
浏览论文内容
中文总结 AI 辅助
本文分析并改进低秩进化策略EGGROLL,通过理论刻画其更新场并引入留一估计器LOO-ROLL,在保持性能的同时减半评估成本,并在后训练中提升准确率。
中文摘要 AI 辅助
EGGROLL通过将密集高斯权重扰动替换为低秩高斯乘积(通常为秩一)使进化策略(ES)对大型语言模型(LLM)变得实用。这种选择在计算上具有吸引力,但在几何上却极为严苛:每个秩一扰动位于环境矩阵空间的零体积子集中,尽管其协方差为单位矩阵。我们刻画了有限秩和非零扰动半径下EGGROLL平均更新场的特征,然后分析了其有限总体估计器的误差。总体场是通过对由扰动平滑的目标函数的梯度应用显式预解式获得的。我们表明,该预解式可能引入非保守分量,并可能反转最优点的局部稳定性。尽管如此,EGGROLL在每一个二次目标上,在任意秩和半径下都是精确的。对于光滑目标,其第一个局部有限秩修正为$O(\sigma^2/r)$,并且非渐近界在光滑性假设下控制了由此产生的场误差。在局部仿射模型下,与密集高斯ES相比,秩一扰动仅将梯度估计器的方差增加了$\frac{2(m+n+1)}{mn+1}$,对于$4096\times4096$矩阵,即$0.098\\%$。然后,我们引入了LOO-ROLL,一种留一估计器,它保留了有限秩总体场,同时将EGGROLL每个方向的两次对偶评估替换为一次。在相同评估成本下,LOO-ROLL在Transformer块中将估计器MSE减半。在十个后训练设置和高达8B参数的模型上,在匹配的墙钟时间下,LOO-ROLL在个体配对测试中改善了七个结果,且没有显著损失。在GSM8K测试集上,0.6B模型的准确率从$38.1\\%$提高到$63.0\\%$,8B模型从$65.9\\%$提高到$80.0\\%$。Transformer测量恢复了预测的有限秩方差,而秩比较显示秩八没有可复现的基于奖励的优势。
英文摘要
EGGROLL (Sarkar et al., 2026) makes evolution strategies (ES) practical for LLMs by replacing dense Gaussian weight perturbations with low-rank Gaussian products, often of rank one. This choice is computationally attractive but geometrically severe: Each rank-one perturbation lies in a zero-volume subset of the ambient matrix space, despite having identity covariance. We characterize the EGGROLL update mean field at finite rank and nonzero perturbation radius as a resolvent applied to the gradient of the perturbation-smoothed objective. This transformation can make the mean field nonconservative and reverse the local stability of an optimum. EGGROLL nevertheless recovers the gradient exactly on quadratic objectives at every rank and radius. In finite populations, the additional sampling variance of rank-one perturbations relative to dense Gaussian ES decays inversely with matrix width under a local affine model, and is only $0.098\%$ at width $4096$. Finally, we introduce LOO-ROLL, a leave-one-out estimator that replaces EGGROLL's two antithetic evaluations per direction by one. At equal evaluation cost, LOO-ROLL halves estimator MSE in transformer blocks. Across fourteen post-training settings up to 14B parameters, matched-time comparisons with EGGROLL yield eleven improvements in individual paired tests and no significant loss. At 1.7B, 8B, and 14B parameters, matched-time gains are $2.9$, $14.1$, and $7.9$ percentage points on GSM8K and $12.2$, $8.4$, and $8.1$ points on MATH-500.
发表机构
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。