AI 中文总结
提出ERAHBO方法,将RL学习结果的均值和方差建模为HP配置的函数,通过自适应重采样提升样本效率,在多种RL算法与环境上优于相关基线。
AI 中文摘要
强化学习(RL)在各类复杂任务中已取得显著成功,但RL结果具有高度随机性,其期望性能与变异性通常依赖超参数(HP)配置。本文提出高效风险感知异方差贝叶斯优化(ERAHBO),该贝叶斯优化方法将学习结果的均值与方差均建模为HP配置的函数;ERAHBO旨在识别能实现高平均回报且降低训练运行变异性的HP配置,并通过自适应重采样而非每个HP的固定预算提升HP优化的样本效率。在多种RL算法与环境上的实证评估显示,ERAHBO总体上优于风险中性与风险感知基线,为风险感知回报提供了更高的样本效率。
英文摘要
Reinforcement learning (RL) has shown remarkable success across a wide range of complex tasks. However, RL outcomes can be highly stochastic, and both expected performance and variability often depend on hyperparameter (HP) configurations. We propose efficient and risk-averse heteroscedastic Bayesian Optimization (ERAHBO), a Bayesian optimization method that models both the mean and variance of learning outcomes as functions of the HP configurations. ERAHBO aims to identify HP configurations that achieve high average return while reducing variability across training runs, and it improves the sample efficiency of the HP optimization via adaptive re-sampling rather than a fixed budget per HP. Empirical evaluations across diverse RL algorithms and environments demonstrate that ERAHBO generally outperforms both risk-neutral and risk-averse baselines, delivering improved sample efficiency for risk-averse returns.
CommentsAccepted at RLC'26