AI 中文总结
研究针对梯度下降和无梯度优化器在大型模型长时间学习中的问题,提出一种元学习算法,通过内、外循环,智能体可修改自身权重偏差,外循环用标准零阶方法处理少量参数,有望加速学习。
AI 中文摘要
梯度下降在大型模型中扩展性良好,但在长时间范围内会变得不稳定。无梯度优化器可扩展到任意时间跨度,但受高维度限制。由于学习在大型模型的长时间尺度上进行,这两种方法都不太可能产生加速学习过程的特性。我们提出一种元学习算法,其中智能体学习修改自身权重和偏差。算法由内循环(智能体对自身进行高维优化)和外循环(对外循环进行低维优化)组成,外循环处理参数少,可用标准零阶方法。
英文摘要
Gradient descent scales well to large models, but becomes unstable over long time horizons. Gradient-free optimizers can scale to arbitrary timespans, but are hobbled by high dimensions. Since learning occurs in large models over long timescales, neither of these approaches is likely to produce traits which can accelerate the learning process. Instead, we propose a meta-learning algorithm in which the agent learns to modify its own weights and biases. Our algorithm consists of an inner loop, wherein the agent performs some high-dimensional optimization upon itself, and an outer loop, wherein we perform some low-dimensional optimization upon the inner loop. Since the outer loop handles very few parameters, standard zeroth-order methods may be used.