arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于近端策略优化(PPO)的卫星轨道优化用于空间碎片规避

Satellite Trajectory Optimization via Proximal Policy Optimization for Space Debris Avoidance

Logan Luna, Juan Ortiz Couder, Raul Alejandro Vargas-Acosta

arXiv 2608.09628首次发表:更新:

发表机构

Georgia Institute of Technology; Embry-Riddle Aeronautical University(佐治亚理工学院; 安柏瑞德航空大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对轨道拥堵下卫星碰撞规避难题,提出基于PPO的强化学习策略,经高保真模拟器训练,在GEO任务中碰撞规避成功率达97.5%,优于传统方法,构建了公开可用框架

AI 中文摘要

规避碰撞系统通常用于避免低地球轨道(LEO)和地球静止轨道(GEO)发生碎片事件,但随着巨型星座的发射导致轨道拥堵加剧,此类事件频率不断上升, conjunction 警报和碰撞风险日益常见。当前常用的手动或基于规则的方法难以适配这种不断恶化的动态环境。为应对这一严峻状况,我们提出一种用于自主碰撞规避的强化学习策略,通过近端策略优化(PPO)训练,并结合开源高保真天体动力学模拟器开展训练与评估。在1000次确定性GEO任务中,我们的智能体实现了97.5%的碰撞规避成功率,优于基于规则的基线(成功率20.7%)和脉冲 delta-v 规划基线(成功率27.5%)等传统控制器。为获得这些结果,我们设计了一款模拟器,利用真实和模拟的碎片训练并评估智能体,模拟包含太阳/月球第三体摄动、依赖燃料的推力及可配置碎片场的牛顿二体动力学;智能体通过课程学习和导向生存、足够预测 miss 距离及 delta-v 节约的塑形奖励进行训练。最后,我们的评估采用完全确定性流程,包括共享种子、每次任务日志和遥测导出。本工作是一个公开可用的框架,详见该 https URL

英文摘要

Collision avoidance systems are commonly used to avoid fragmentation events occurring in Low-Earth Orbit (LEO) and Geosynchronous Equatorial Orbit (GEO). However, these events have been growing in frequency as orbital congestion worsens with the launch of megaconstellations. Consequently, conjunction alerts and collision risks are becoming increasingly common. Current practices, which are commonly manual or rule-based, have difficulty scaling to these worsening dynamic environments. To address this intensifying situation, we propose a reinforcement-learning policy for autonomous collision avoidance, trained via Proximal Policy Optimization (PPO) along with an open-source, high-fidelity astrodynamics simulator for training and evaluation. In 1,000 deterministic GEO episodes, our agent achieves a 97.5% collision avoidance success rate, outperforming traditional controllers such as a rule-based baseline (20.7% success) and an impulsive delta-v planner baseline (27.5% success). To achieve these results, we designed a simulator to train and evaluate our agent, using real-world and simulated debris. We simulate Newtonian two-body dynamics using Sun/Moon third-body perturbations, fuel-dependent thrust, and configurable debris fields. The agent is trained with curriculum learning and shaped rewards oriented toward encouraging survival, adequate projected miss distance, and delta-v conservation. Finally, our evaluation consisted of a fully deterministic pipeline, including shared seeds, per-episode logs, and telemetry exports. Our work is a publicly available framework at https://purl.org/sat-trajectory-avoidance

Comments18 pages, 16 figures. Published in IEEE Access, vol. 14, pp. 18138-18154, 2026

Journal refL. Luna, J. Ortiz Couder, and R. A. Vargas-Acosta, "Satellite Trajectory Optimization via Proximal Policy Optimization for Space Debris Avoidance," IEEE Access, vol. 14, pp. 18138-18154, 2026

DOI:10.1109/ACCESS.2026.3655237

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑