arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

矩阵函数的无梯度优化

Gradient-Free Optimization for Matrix functions

Sawyer Allen, Cash Cherry, Aidan Eck, Stephen Becker, Daniel McKenzie

arXiv 2609.03170首次发表:更新:

发表机构

Department of Mathematics, UCSB; Department of Applied Mathematics, CU Boulder; Department of Applied Mathematics and Statistics, Colorado School of Mines(加州大学圣塔芭芭拉分校数学系; 科罗拉多大学博尔德分校应用数学系; 科罗拉多矿业学院应用数学与统计系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对仅能获取方向导数的矩阵函数优化场景,提出替代随机梯度估计器的方案,结合矩阵感知技术与谱下降优化器,通过两项数值实验验证其可加速收敛至良好近似解。

AI 中文摘要

我们考虑仅能获取方向导数而非完整梯度时,对矩阵变量的光滑(可能非凸)函数进行优化的任务。该场景出现在消费级硬件上微调大型神经网络时:网络权重为矩阵,内存限制排除了反向模式自动微分,但正向模式仍可提供方向导数。我们将该场景下的梯度估计视为信号处理领域的结构化恢复问题。从这一视角出发,我们做出三项贡献:其一,我们提出了标准随机梯度估计器的替代方案,其差异在于将采样算子的伴随矩阵替换为其伪逆;其二,当梯度满足近似低秩条件时,来自矩阵感知的技术可生成一系列高精度梯度估计器,且这些估计器可嵌入任意一阶方法;其三,我们指出这类估计器的计算成本虽高,但可通过与谱下降(spectral descent)等感知矩阵的优化器结合来分摊成本。具体而言,梯度估计器会计算梯度的分解,使谱下降的投影步骤可在无额外成本的情况下完成。我们针对具有近似低秩梯度的合成函数开展两项精心设计的数值实验,验证了上述结论,结果表明,通过利用低秩特性,可更快收敛至良好的近似解。

英文摘要

We consider the task of optimizing smooth, possibly non-convex functions of a matrix variable given access only to directional derivatives rather than full gradients. This setting arises when fine-tuning large neural networks on consumer-grade hardware: the network's weights are matrices, memory constraints rule out backward-mode automatic differentiation, but directional derivatives remain available through forward mode. We frame gradient estimation in this setting as a structured recovery problem, in the spirit of signal processing. From this perspective we provide three contributions. First, we introduce an alternative to the standard random gradient estimator; the difference corresponds to replacing the adjoint of the sampling operator with its pseudoinverse. Second, when the gradient satisfies an approximate low-rank condition, techniques from matrix sensing yield a family of highly accurate gradient estimators that drop into any first-order method. Third, we note that while the computational cost of such estimators is high, this can be amortized by combining them with a matrix-aware optimizer such as spectral descent. Specifically, the gradient estimator computes a factorization of the gradient, allowing for the projection step of spectral descent to be done at no extra cost. We demonstrate our findings with two careful numerical experiments on synthetic functions with approximately low-rank gradients. We show that by exploiting this low-rank property one obtains much faster convergence to good approximate solutions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑