aRieL - 用于Ariel太空望远镜调度的强化学习
aRieL - Reinforcement Learning for the Ariel Space Telescope Scheduling
- University College London(伦敦大学学院)
- INAF-Osservatorio Astronomico di Palermo(意大利国家天体物理研究所巴勒莫天文台)
- ML Analytics(ML分析公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对Ariel太空望远镜的复杂调度问题,提出基于强化学习的aRieL框架,结合Set Transformer PPO智能体,在多个巡天目标间取得强折中,并展现出良好的泛化能力。
AI中文摘要:
Ariel任务将对数百颗系外行星大气进行群体级巡天,这要求在一个庞大且多样化的候选目标目录中安排观测。通过将Ariel调度问题构建为一个长时域优化问题,其中个体决策既影响任务时间,也影响巡天后期可用的科学机会,强化学习(RL)可作为目标选择的框架。我们引入了aRieL,一个模拟观测环境,结合了Set Transformer近端策略优化(PPO)智能体,训练其生成调度策略。与一组基线启发式策略相比,该RL策略始终能在竞争性巡天目标之间找到强折中,产生大型Tier 1样本,同时保持高Tier 3完成率、群体覆盖率和观测效率。当候选目录大幅扩展时,该行为得以保留,而奖励函数的修改会导致学习到的观测策略发生相应变化。我们进一步表明,在基线任务场景上训练的策略可以在无需重新训练的情况下泛化到需要重复Tier 3观测的修改场景,揭示了与更广泛巡天之间的权衡。这些结果表明,强化学习为大规模天文调度提供了一种灵活的方法,其中观测策略可以直接响应任务状态和科学优先级的变化,并促使其在具有类似复杂、状态依赖调度问题的观测站中得到更广泛的探索。
英文摘要:
The Ariel mission will conduct a population level survey of hundreds of exoplanet atmospheres, requiring observations to be scheduled across a large and diverse catalogue of potential targets. By framing Ariel scheduling as a long-horizon optimisation problem in which individual decisions affect both mission time and the scientific opportunities available later in the survey, Reinforcement Learning (RL) can be utilised as a framework for target selection. We introduce aRieL, a simulated observing environment coupled to a Set Transformer Proximal Policy Optimisation (PPO) agent that is trained to generate a scheduling policy. The RL policy consistently finds a strong compromise between competing survey objectives, producing large Tier~1 samples while maintaining high Tier~3 completion, population coverage, and observing efficiency when compared to a set of baseline heuristic polices. This behaviour is retained when the candidate catalogue is substantially expanded, while modifications to the reward function produce corresponding changes in the learned observing strategy. We further show that a policy trained on the baseline mission scenario can generalise without retraining to a modified scenario requiring repeated Tier~3 observations, revealing the resulting trade-off with the wider survey. These results demonstrate that reinforcement learning provides a flexible approach to large-scale astronomical scheduling, in which the observing strategy can respond directly to changes in mission state and to the scientific priorities, and motivate its wider exploration for observatories with similarly complex, state dependent scheduling problems.