arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

低轨卫星网络中用于双层空中在线联邦学习的自适应波束跳变与功率控制

Adaptive Beam Hopping and Power Control for Dual-Layer Over-the-Air Online Federated Learning in LEO Satellite Networks

Zhendong Li, Shaojie Wang, Zhou Su, Zihao Zhang, Haixia Peng, Nan Cheng, Ying Wang, Wen Chen

arXiv 2609.03202首次发表:更新:

发表机构

Xidian University; Beijing University of Posts and Telecommunications; Shanghai Jiao Tong University(西安电子科技大学; 北京邮电大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对低轨卫星网络双层空中在线联邦学习的长期数据利用率最大化问题,提出基于PPO的深度强化学习框架,联合优化自适应波束跳变与功率控制,性能优于基准方案。

AI 中文摘要

本文研究低轨(LEO)卫星网络中空中(OTA)计算支持的在线联邦学习(FL)。具体而言,我们考虑一种双层OTA聚合架构:地面设备通过上行OTA聚合向服务卫星上传模拟模型更新,卫星再通过第二轮OTA聚合将聚合信号转发至数据处理中心。我们构建了一个长期数据利用率最大化问题,其中设备持续收集新数据,未训练样本的新鲜度逐渐降低。该问题受卫星波束预算、发射功率限制以及控制端到端聚合失真的全局均方误差(MSE)约束,由此产生一个耦合混合整数非线性规划(MINLP)问题,涉及紧密耦合的离散波束跳变决策与连续功率控制。由于组合动作空间和非凸约束,该问题为NP难问题,计算上难以处理。此外,时变的卫星拓扑和动态数据生成使其成为序列决策问题,需要自适应在线调度。为解决这些问题,我们将该问题建模为马尔可夫决策过程,并开发了一个基于近端策略优化(PPO)的深度强化学习框架,该框架利用感知MSE的奖励函数平衡数据利用率与聚合精度,联合优化自适应波束跳变与功率控制。数值仿真结果验证,所提算法在满足MSE要求的同时,始终优于其他基准方案,实现了更优的长期数据利用率和更快的FL收敛速度。

英文摘要

This paper investigates over-the-air (OTA) computation enabled online federated learning (FL) in low-Earth orbit (LEO) satellite networks. Specifically, we consider a dual-layer OTA aggregation architecture, where ground devices upload analog model updates to serving satellites via uplink OTA aggregation, and satellites forward the aggregated signals to a data processing center through the second round OTA aggregation. Then, we formulate a long-term data-utilization maximization problem in which devices continuously collect new data and untrained samples gradually lose freshness. The problem is subject to the satellite beam budget, transmit-power limit, and global mean squared error (MSE) constraint that governs end-to-end aggregation distortion. This yields a coupled mixed-integer nonlinear programming (MINLP) problem, involving tightly coupled discrete beam-hopping decisions and continuous power control. Due to the combinatorial action space and nonconvex constraints, the problem is NP-hard and computationally intractable. Furthermore, the time-varying satellite topology and dynamic data generation render it a sequential decision-making problem, necessitating adaptive online scheduling. To address these issues, we cast the problem as a Markov decision process and develop a proximal policy optimization (PPO)-based deep reinforcement learning framework that jointly optimizes adaptive beam hopping and power control, using an MSE-aware reward to balance data utilization and aggregation accuracy. Numerical simulation results verify that the proposed algorithm consistently outperforms other benchmark schemes, achieving superior long-term data utilization and faster FL convergence while satisfying the MSE requirement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑