arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习协调进化搜索:优化与结构模型更新中面向动态DE-CMA-ES协调的进度感知深度强化学习

Learning to Orchestrate Evolutionary Search: Progression-Aware Deep Reinforcement Learning for Dynamic DE-CMA-ES Coordination in Optimization and Structural Model Updating

Lechen Li, Rongye Shi, Wanhuan Zhou

arXiv 2610.11546首次发表:更新:

发表机构

State Key Laboratory of Internet of Things for Smart City, University of Macau; College of Water Conservancy and Hydropower Engineering, Hohai University; School of Artificial Intelligence, Beihang University(澳门大学物联网智慧城市建设国家重点实验室; 河海大学水利水电学院; 北京航空航天大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出DRL-DCO算法,通过DDPG智能体动态协调DE与CMA-ES的搜索过程,在高维优化及结构模型更新任务上取得了优于现有算法的收敛精度与鲁棒性。

AI 中文摘要

求解高维结构模型更新问题需要一种能够在具有相关参数的复杂非凸空间中导航的算法。现有混合进化算法通常依赖静态架构或固定切换规则,导致搜索阶段脱节。为解决该问题,本研究提出一种由深度强化学习控制的动态DE-CMAES协调算法(DRL-DCO),其中基于深度确定性策略梯度(DDPG)的演员-评论家智能体将进化过程作为单一统一系统持续控制,而非算法的机械拼接。在进度感知状态表示和多样性感知奖励的引导下,该智能体在差分进化(DE)基于差向量的探索与协方差矩阵自适应进化策略(CMA-ES)基于协方差的利用之间灵活重新分配计算资源,同时共同调控种群规模、精英保留和重启机制以逃离局部最优。这使得DRL-DCO能够在各代中自主在探索主导、利用主导和混合策略模式之间切换。在训练阶段之外,训练后的演员可在无监督推理模式下运行,其中内化的策略仅通过前向推理从观测到的搜索状态自主协调DE和CMA-ES的控制,无需评论家评估或权重更新,从而在保持完全有效性的同时实现更快部署。在高维单目标优化基准和IASC-ASCE结构健康监测基准上的验证表明,DRL-DCO与最先进的自适应及混合进化算法、以及单算子DRL控制基线相比,实现了更优的收敛精度和鲁棒性。

英文摘要

Solving high-dimensional structural model updating problems requires an algorithm capable of navigating complex, non-convex landscapes with correlated parameters. Existing hybrid evolutionary algorithms typically rely on static architectures or fixed switching rules, resulting in disjointed search phases. To address this, this study proposes a Deep Reinforcement Learning-governed dynamic DE-CMAES Orchestration (DRL-DCO) algorithm, in which a Deep Deterministic Policy Gradient (DDPG)-based actor-critic agent continuously governs the evolutionary process as a single, unified system rather than a mechanical concatenation of algorithms. Guided by a progression-aware state representation and a diversity-informed reward, the agent fluidly reallocates computational resources between the differencevector-based exploration of Differential Evolution (DE) and the covariance-guided exploitation of CMA-ES, while jointly regulating population size, elite preservation, and a restart mechanism to escape local optima. This allows DRL-DCO to autonomously transition between exploration-dominant, exploitation-dominant, and mixed-strategy regimes across generations. Beyond the training phase, the trained actor can operate in a supervision-free inference mode, where the internalized policy autonomously orchestrates DE and CMA-ES control from observed search states through forward inference alone, without critic evaluation or weight updates, enabling faster deployment while retaining full effectiveness. Validated on high-dimensional single-objective optimization benchmarks and the IASC-ASCE structural health monitoring benchmark, DRL-DCO achieves superior convergence accuracy and robustness compared to state-of-the-art adaptive and hybrid evolutionary algorithms, as well as single-operator DRL-governed baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑