可达性信息增强的多脉冲行星际转移强化学习
Reachability-Informed Reinforcement Learning for Multi-Impulse Interplanetary Transfers
浏览论文内容
中文总结 AI 辅助
本文提出可达性分析增强的强化学习(RARL),通过将中间航路点选择与兰伯特重构结合,实现多脉冲行星际转移的轨迹设计,在基准上达到接近最优的机动成本并显著提升策略跨发射条件的复用可行性。
中文摘要 AI 辅助
强化学习为航天器轨迹设计提供了一种可复用的序贯决策机制的前景,这促使人们设计将学习到的决策与底层机动几何结构相连接的策略接口。本文针对确定性多脉冲行星际转移问题,提出了可达性分析增强的强化学习(RARL)方法,将中间航路点选择置于学习决策过程的核心。局部一阶可达性图将有限的速度脉冲映射为下一节点位置的椭球集合,策略在该集合内选择其航路点。随后,通过兰伯特重构确定相应的机动,以沿动力学一致的弹道弧段到达所选航路点,从而将学习到的转移几何选择与经典天体动力学相结合。终端双脉冲重构完成交会,并辅以用于奖励塑形的线性机动需求评估。数值研究在二体地球-火星基准问题上刻画了该接口的性能。在三次独立训练运行中,RARL实现了平均机动成本10.23 km/s,比经过验证的局部序贯凸规划参考值高出1.72%。在分散初始状态上的训练将策略复用扩展到具有固定目标状态和转移时长的发射族。三个独立训练的多状态策略均完成了全部10,000次留出蒙特卡洛发射,且无脉冲上限违规,而单状态策略的平均可行性率仅为6.49%。这种更广泛的采样可行性伴随着平均名义机动成本增加0.61%,且无需在发射之间进行额外训练。这些结果表明,可达性信息增强的决策接口支持基准质量的轨迹构建以及跨分散发射条件的策略复用。
英文摘要
Reinforcement learning offers the prospect of a reusable sequential decision-making mechanism for spacecraft trajectory design, motivating policy interfaces that connect learned decisions to the underlying maneuver geometry. This paper develops Reachability Analysis-Informed Reinforcement Learning (RARL) for deterministic multi-impulse interplanetary transfers, placing intermediate waypoint selection at the center of the learned decision process. Local first-order reachability maps bounded velocity perturbations into an ellipsoidal set of next-node positions, within which the policy selects its waypoint. Lambert reconstruction then determines the corresponding maneuver to reach this selected waypoint along a dynamically consistent ballistic arc, coupling learned transfer-geometry selection with classical astrodynamics. A terminal two-impulse reconstruction completes the rendezvous, supported by a linear maneuver-demand assessment used for reward shaping. Numerical studies characterize this interface on a two-body Earth-Mars benchmark. Across three independent training runs, RARL achieves a mean maneuver cost of 10.23 km/s, 1.72% above a validated local sequential convex programming reference. Training over dispersed initial states extends policy reuse across a departure family with fixed target state and transfer duration. Each of the three independently trained multi-state policies completes all 10,000 held-out Monte Carlo departures without impulse-cap violations, compared with a mean feasibility rate of 6.49% for single-state policies. This broader sampled feasibility is accompanied by a 0.61% increase in mean nominal maneuver cost, without further training across departures. These results demonstrate that a reachability-informed decision interface supports benchmark-quality trajectory construction and policy reuse across dispersed departure conditions.
发表机构
- Waipapa Taumata Rau– University of Auckland(怀帕帕陶马塔劳——奥克兰大学)
- European Space Agency(欧洲空间局)
机构由 AI 辅助整理,请以论文原文为准。