arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不确定性感知的在线算法选择:基于模拟器集成方法

Uncertainty-Aware Selection of Online Algorithms with Simulator Ensembles

Yongyi Guo, Zifan Xu, Ziping Xu, Kelly W. Zhang

arXiv 2609.32170首次发表:更新:

AI 中文总结

本文提出不确定性感知的模拟器集成选择方法,以缓解离线数据有限时Plug-In选择的不稳定性,并在理论上和机器人控制实验中证明其能提升在线强化学习的可靠性与性能。

AI 中文摘要

在线强化学习的性能关键依赖于设计选择,尤其是那些影响探索的选择。这些选择通常通过将模拟器拟合到离线数据、在该模拟器中评估候选算法,并部署表现最佳的一种来进行。最简单的Plug-In选择规则仅选择在拟合模拟器上表现最佳的算法,当用于拟合模拟器的离线数据有限时,这种评估变得不可靠。我们研究了不确定性感知选择方法,该方法构建一个模拟器集成——例如,通过自助重采样获得——并选择在集成中平均性能最佳的在线算法。虽然基于集成的方法已被用于缓解分布偏移和促进模拟到现实的迁移,我们正式证明这种方法在拟合模拟器时可以缓解有限数据的影响,并且在理论上与Plug-In选择相比,在多臂老虎机中具有显著的遗憾减少。我们还在深度强化学习实验中实证研究了不确定性感知选择方法,涉及选择奖励塑形超参数的机器人控制任务,并表明它导致更可靠的选择和改进的在线性能。

英文摘要

The performance of online reinforcement learning depends critically on design choices, especially those that affect exploration. These choices are often selected by fitting a simulator to offline data, evaluating candidate algorithms in that simulator, and deploying the best-performing one. The simplest Plug-In selection rule simply selects the best performing algorithm on the fitted simulator, making evaluations unreliable when the offline data used to fit the simulator are limited. We investigate Uncertainty-Aware selection, which forms an ensemble of simulators---for example, obtained by bootstrap resampling---and selects the online algorithm with the best average performance across the ensemble. While ensemble-based approaches have been used to mitigate distribution shift and facilitate sim-to-real transfer, we formally show that this approach can mitigate the effects of limited data when fitting the simulator and theoretically has significant regret gains compared to Plug-In selection in multi-armed bandits. We also empirically investigate the Uncertainty-Aware selection approach in deep RL experiments on robotic control tasks that involve selecting reward-shaping hyperparameters, and show that it leads to more reliable selection and improved online performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑