AI 中文总结
针对卫星辅助UAM网络的多子带调度问题,提出几何感知集注意力PPO方法GeoSetPPO,结合两阶段训练策略,实现了更低调度延迟与更优性能。
AI 中文摘要
本文研究空天地一体化协同网络中城市空中交通(UAM)的下行链路调度问题:多个地面站(GS)采用窄三维波束并在多子带间共享频谱,卫星提供正交频段服务选项;时变的几何结构与定向干扰要求联合决策基站关联、GS子带分配及发射功率。我们构建有限时域混合离散-连续问题,仅利用UAM的位置与速度信息,在最大化总速率的同时惩罚切换次数与GS过载。为解决组合调度问题,提出GeoSetPPO,即几何感知集注意力近端策略优化(PPO)方法,该方法输出每个UAM的离散关联与子带决策,采用排列不变表示;在给定调度方案下,GS的功率由逐时隙连续凸近似(SCA)模块计算,需满足每个GS的功率预算及最小信干噪比(SINR)约束。为降低训练成本并提升稳定性,采用两阶段训练策略,将奖励评估从均匀功率分配过渡到基于SCA的功率分配。仿真结果显示,所提方法在考虑的训练设置下收敛稳定,回报高于基于多层感知器(MLP)和Transformer的PPO;与基于算法及距离的调度器相比,奖励和调度可行性表现更优;在更大规模网络中,相较于先前基于算法的方法,GeoSetPPO还将调度延迟从40.84 ms降至2.90 ms。
英文摘要
In this paper, we investigate downlink scheduling for urban air mobility (UAM) in a cooperative space-air-ground integrated network. Multiple ground stations (GSs) employ narrow three-dimensional beams and share spectrum across multiple subbands, while a satellite provides an orthogonal-band service option. Rapidly time-varying geometry and directional interference require joint decisions on base station association, GS subband assignment, and transmit powers. We formulate a finite-horizon mixed discrete-continuous problem that maximizes sum rate while penalizing handovers and GS overload, using only UAM positions and velocities. To address the combinatorial scheduling problem, we propose GeoSetPPO, a geometry-aware set-attention proximal policy optimization (PPO) method that outputs per-UAM discrete association and subband decisions with permutation-invariant representations. Conditioned on each schedule, GS powers are computed by a per-slot successive convex approximation (SCA) module under per-GS power budgets and minimum signal-to-interference-plus-noise ratio (SINR) constraints. To reduce training cost and improve stability, we adopt a two-stage training strategy that transitions reward evaluation from uniform power to SCA-based power allocation. Simulations demonstrate stable convergence, higher returns than multi-layer perceptron (MLP)- and Transformer-based PPO under the considered training setting, and favorable reward and schedule-feasibility performance relative to algorithm-based and distance-based schedulers. In the larger evaluated network, GeoSetPPO also reduces the scheduling latency from 40.84 ms to 2.90 ms relative to the previous algorithm-based method.
Comments16 pages, 14 figures