AI 中文总结
针对多轨道卫星网络的联合切换管理与功率分配问题,提出结合MAPPO与TarMAC的多智能体强化学习策略,在保证接近贪心方案吞吐量的同时大幅减少切换次数,性能优于仅低轨方案与保守停留策略。
AI 中文摘要
未来第六代非地面网络有望将低轨(LEO)、中轨(MEO)和地球静止轨道(GEO)卫星相结合,其互补层必须在低轨卫星快速动态变化下,通过联合用户关联、功率分配和切换管理进行协调。本文将该问题建模为混合整数非线性规划,并将其分解为:多智能体强化学习(MARL)策略用于选择关联关系,以及每个时隙精确求解的凸功率分配子问题,该子问题定义了MARL部分的奖励。关联策略采用多智能体近端策略优化(MAPPO)和目标多智能体通信(TarMAC)机制进行训练,并通过编码依赖轨道层切换惩罚的状态,使其感知轨道层。在基于肯尼亚内罗毕真实两行元数据构建的现实多星座场景下评估,所提策略达到了贪心信噪比(SNR)最大化方案92%的吞吐量,同时触发的切换次数减少4倍以上,且相比保守停留启发式算法,吞吐量提升约14%。与相同架构的仅低轨学习策略相比,该策略通过将部分用户卸载至MEO和GEO层,实现了略高的吞吐量和更少的切换,这种涌现的多轨道行为促成了其良好的吞吐量与切换权衡。
英文摘要
Future sixth-generation non-terrestrial networks are expected to combine low Earth orbit (LEO), medium Earth orbit (MEO), and geostationary Earth orbit (GEO) satellites, whose complementary layers must be coordinated through joint user association, power allocation, and handover management under fast LEO dynamics. This paper studies this problem by formulating it as a mixed-integer nonlinear program and decomposing it into a multi-agent reinforcement learning (MARL) policy that selects the associations and a convex power-allocation subproblem solved exactly at each time slot that defines the reward of the MARL part. The association policy is trained with multi-agent proximal policy optimization (MAPPO) and the targeted multi-agent communication (TarMAC) mechanism, and is made aware of the orbital layer through a state that encodes layer-dependent handover penalties. Evaluated on a realistic multi-constellation scenario built from real two-line element data over Nairobi, Kenya, the proposed policy reaches 92% of the throughput of a greedy signal-to-noise ratio (SNR) maximizing scheme while triggering more than four times fewer handovers, and improves throughput by roughly 14% over a conservative stay heuristic. Compared to an LEO-only learned policy of identical architecture, it attains slightly higher throughput with fewer handovers by offloading a fraction of the users to the MEO and GEO layers, an emergent multi-orbit behavior that drives its favorable throughput and handover trade-off.