arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于强化学习超启发式算法的运行时分析

On the Runtime Analysis of Reinforcement Learning Hyper-Heuristics

Pietro S. Oliveto, Zhenyu Wang, Peizhou Wu, Mengqing Xu

arXiv 2607.22036首次发表:更新:

AI 中文总结

研究强化学习超启发式算法(RLHH),通过证明配备特定随机局部搜索算子的RLHH在适当参数值下能在最佳预期运行时优化LeadingOnes基准函数,且实验显示其在实际问题规模中比广义随机梯度HH更快。

AI 中文摘要

选择超启发式算法(HHs)通过从一组低级启发式算法中选择在优化过程的每个阶段应用哪一个来实现算法设计自动化。最近关于选择超启发式算法对标准基准函数的性能有一些严格证明的令人印象深刻的结果。然而,与现实世界应用中通常使用的机器学习技术相比,这些HHs采用的学习机制相当简化。本文分析了文献中的一种强化学习超启发式算法(RLHH)。之前唯一可用的结果证明,对于广泛的参数设置,RLHH不能学会为标准的LeadingOnes基准函数适当地选择启发式算法。本文严格证明,配备两个随机局部搜索算子RLS_1和RLS_2的RLHH,在适当的参数值下,能在两个算子可达到的最佳预期运行时内优化LeadingOnes基准函数,直至低阶项。实验表明,对于实际问题规模,它比之前被证明也具有直至低阶项的最优预期运行时的广义随机梯度HH更快。

英文摘要

Selection Hyper-heuristics (HHs) automate algorithmic design by selecting from a set of low-level heuristics which one to apply at each stage of the optimisation process. Several impressive results have been recently rigorously proven regarding the performance of selection hyper-heuristics (HHs) for standard benchmark functions. However, the learning mechanisms employed by these HHs are considerably simplified compared to the machine learning techniques typically used in real world applications. In this paper we analyse a Reinforcement Learning Hyper-heuristic (RLHH) from the literature. The only previous result available proved that for a wide range of parameter settings, RLHH does not learn to select heuristics appropriately for the standard LeadingOnes benchmark function. In this paper, we rigorously prove that with appropriate parameter values RLHH equipped with two random local search operators, RLS_1 and RLS_2 optimises the LeadingOnes benchmark function in the best possible expected runtime achievable with the two operators up to lower order terms. Experiments show that for realistic problem sizes it is faster than the Generalised Random Gradient HH which was previously proven to also have optimal expected runtime up to lower order terms.

CommentsPublished in PPSN 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑