强化学习引导的进化策略优化用于偏好可调的异构敏捷地球观测卫星调度
Reinforcement Learning-Guided Evolutionary Policy Optimization for Preference-Adjustable Heterogeneous Agile Earth Observation Satellite Scheduling
浏览论文内容
中文总结 AI 辅助
针对异构敏捷地球观测卫星调度难题,本文提出强化学习辅助的算子选择模因进化算法,在有限预算下实现更高加权效用与稳定收敛,验证了方法鲁棒性与算子选择的有效性。
中文摘要 AI 辅助
异构敏捷地球观测卫星(AEOS)调度需要在卫星依赖的可见窗口、姿态机动要求、能耗及星上存储约束下完成任务选择、卫星分配与观测序列规划。由于卫星在轨道访问、机动能力及载荷资源上存在差异,同一任务在不同平台上可能有不同的可行窗口、过渡成本及资源消耗模式,这增加了统一建模与高效优化的难度。为解决该问题,本文提出一种用于异构AEOS调度的进化策略优化框架,该框架支持偏好可调的加权目标。在建模层,将基于分配的间接编码与基于解码器的等效成本评估相结合,在保留卫星依赖约束的同时,将任务收益、节能及载荷平衡整合为可解释的标量效用。在优化层,将调度解码、基于种群的搜索及在线演员-评论家算子控制解耦,使强化学习选择高层搜索算子而非直接构建调度。基于该框架,开发了强化学习辅助的算子选择模因进化算法(RLOSMEA),以在有限的函数评估预算下协调全局探索、可行性恢复及局部优化。在不同异构AEOS场景上的实验表明,RLOSMEA相较于代表性元启发式基线算法,能实现更高的整体加权效用与更稳定的收敛性。敏感性分析与学习行为分析进一步证实了所提方法的鲁棒性及强化学习引导的算子选择的有效性。
英文摘要
Heterogeneous agile Earth observation satellite (AEOS) scheduling requires task selection, satellite assignment, and observation sequencing under satellite-dependent visibility windows, attitude maneuvering requirements, energy consumption, and onboard storage constraints. Since satellites differ in orbital access, maneuvering capability, and payload resources, the same task may have different feasible windows, transition costs, and resource-consumption patterns on different platforms, which increases the difficulty of unified modeling and efficient optimization. To address this problem, this paper proposes an evolutionary policy optimization framework for heterogeneous AEOS scheduling with preference-adjustable weighted objectives. In the modeling layer, assignment-based indirect encoding is combined with decoder-based equivalent-cost evaluation to retain satellite-dependent constraints while integrating task gain, energy saving, and load balance into an interpretable scalar utility. In the optimization layer, schedule decoding, population-based search, and online actor-critic operator control are decoupled, so that reinforcement learning selects high-level search operators rather than constructing schedules directly. Based on this framework, a reinforcement-learning-assisted operator-selection memetic evolutionary algorithm (RLOSMEA) is developed to coordinate global exploration, feasibility recovery, and local refinement under a limited function-evaluation budget. Experiments on different heterogeneous AEOS scenarios show that RLOSMEA achieves higher overall weighted utility and more stable convergence than representative metaheuristic baselines. Sensitivity and learning-behavior analyses further confirm the robustness of the proposed method and the effectiveness of reinforcement-learning-guided operator selection.
发表机构
- College of Intelligent Science and Engineering, Harbin Engineering University(哈尔滨工程大学智能科学与工程学院)
- School of Information Science and Technology, Dalian Maritime University(大连海事大学信息科学与技术学院)
- Silesian University of Technology(西里西亚工业大学)
- University of Alberta(阿尔伯塔大学)
- Constructor University(康斯特鲁克托大学)
- Istinye University(伊斯蒂涅大学)
机构由 AI 辅助整理,请以论文原文为准。