arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24471cs.AI

基于隐式Q学习引导的蚁群优化算法的敏捷卫星海上移动目标观测调度

Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target observation scheduling with agile satellites

  • Harbin Engineering University(哈尔滨工程大学)
  • North University of China(中北大学)
  • Dalian Maritime University(大连海事大学)

机构由 AI 辅助整理,请以论文原文为准。

He Wang, Junyu Wu, Yeye Liu, Yifan Zhou, Jie Zhang, Hui Li, Yanjie Song, Liang Li

AI总结:

针对敏捷卫星海上移动目标观测调度难题,提出IQACO算法,将隐式Q学习嵌入蚁群优化以自适应调整参数,实验显示其观测效益更高、收敛更快且稳定性好。

AI中文摘要:

敏捷地球观测卫星的海上移动目标观测调度是一个动态、序列依赖的组合优化问题。海面目标持续移动,导致可行观测窗口随目标运动和卫星轨道几何形状变化。调度器必须在时间窗口、姿态机动、星载资源和云影响可用性约束下,联合确定任务选择、卫星分配、观测窗口选择和观测排序。本文提出一种名为IQACO的隐式Q学习引导的蚁群优化方法,用于多卫星海上移动目标观测调度。IQACO并非直接学习任务选择策略,而是将离线隐式Q学习模块嵌入构造性蚁群优化,以自适应调整信息素因子、启发式因子和蒸发率。紧凑的搜索状态表示捕获信息素分布、当前及历史最优解质量和迭代进度。在线调度时,蚁群优化构造可行观测序列,同时学习的策略根据当前搜索状态调节探索与利用。在14种不同规模和卫星配置的场景实验显示,IQACO在所有场景中均获得最高平均观测效益,较传统蚁群优化算法提升3.40%至9.40%,收敛速度更快,且在不同目标权重设置下保持稳定。这些结果表明,离线价值学习为受约束的海上移动目标观测调度提供了有效的自适应搜索控制机制。

英文摘要:

Maritime moving-target observation scheduling with agile Earth observation satellites is a dynamic, sequence-dependent combinatorial optimization problem. Sea-surface targets move continuously, causing feasible observation windows to vary with target motion and satellite orbital geometry. The scheduler must jointly determine task selection, satellite assignment, observation-window selection, and observation ordering under time-window, attitude-maneuvering, and onboard-resource constraints. This paper proposes an implicit Q-learning-bootstrapped ant colony optimization method, termed IQACO, for multi-satellite maritime moving-target observation scheduling. Rather than directly learning a task-selection policy, IQACO embeds an offline implicit Q-learning module into constructive ant colony optimization to adaptively adjust the pheromone factor, heuristic factor, and evaporation rate. A compact search-state representation captures pheromone distribution, current and historical-best solution quality, and iteration progress. During online scheduling, ant colony optimization constructs feasible observation sequences, while the learned policy adjusts the search behavior according to the current search state. Experiments on 14 scenarios with different scales and satellite configurations show that IQACO consistently outperforms the compared algorithms, improving the mean objective value over conventional ant colony optimization by 2.86%-9.41%. Further comparative and supplementary experiments demonstrate its effectiveness and robustness across different scheduling conditions and problem settings. These results indicate that offline value learning provides an effective adaptive search-control mechanism for constrained maritime moving-target observation scheduling.

补充信息

↑