用于选择优化器的强化学习
Reinforcement learning to choose optimizers
- Delft University of Technology(代尔夫特理工大学)
- Brown University(布朗大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出将优化器选择建模为序列决策问题的强化学习方法,在未见过的问题上除最小预算外均优于组合中所有优化器,且分布偏移下保持稳健性。
AI中文摘要:
不存在对所有问题都最优的单一优化方法,且最合适的优化器选择会在运行过程中发生变化。现有在执行过程中更换优化器的方法通常会预先确定部分策略:组合(portfolio)被限制为一类算法,切换在固定时间进行一次,或决策频率被视为超参数而非学习得到的参数。我们提出“用于选择优化器的强化学习”,将优化算法选择建模为序列决策问题。在每次决策时,循环策略会读取当前运行状态,决定接下来应使用哪个优化器以及使用时长。组合包含基于梯度和无导数的优化器,每次切换会传递当前最优解和代表性步长。上下文代理(context proxy)会对专家头(expert heads)上的门控网络进行条件设置,训练采用解耦的演员-评论家(decoupled actor-critic)算法,其回报使用评估时所用的相同经验运行时间分布指标表示。训练任务与组合被共同设计,以避免单一优化器占据主导。在未见过的问题上,学习到的策略除了最小预算外,在所有预算下都优于组合中的每个优化器,且在分布偏移下保持稳健性。
英文摘要:
No single optimization method is uniformly best for all problems, and the most suitable optimizer choice can change during a run. Existing approaches that change optimizer during execution typically predetermine part of the strategy: the portfolio is restricted to one algorithm class, the switch occurs once at a fixed time, or the frequency of decisions is treated as a hyperparameter rather than a learned one. We introduce "Reinforcement Learning to Choose Optimizers", which formulates the optimization algorithm choice as a sequential decision-making problem. At each decision, a recurrent policy reads the current run state and decides both which optimizer should be used next and for how long. The portfolio includes both gradient-based and derivative-free optimizers, and each switch passes on the current best solution and a representative step size. A context proxy conditions a gating network over expert heads, and training employs a decoupled actor-critic whose return is expressed in the same empirical runtime distribution metric used at evaluation. Training tasks and portfolio are designed jointly so that no optimizer dominates. On unseen problems, the learned policy outperforms every portfolio optimizer at all but the smallest budgets, and it remains robust under distribution shift.